Quiz solutions

Answers

  1. B — the “small, read-only, broadcast to many threads” pattern is what constant memory is for.

  2. A — cudaMallocManaged / managed needs no explicit copy. B works but needs two explicit copies. C cannot work: host code cannot dereference a cudaMalloc pointer, so there is no “host side” to do arithmetic on. D leaves the two buffers unrelated — the device buffer is never filled and the result never comes back (and plain malloc memory is not accessible to kernels without unified memory).

  3. B — automatic page migration is the defining feature of managed memory.

  4. C — “managed” does not imply “synchronised”.

  5. A — the canonical workflow of explicit (manual) memory management.

  6. A — the argument order (destination first) and the unit of count (bytes, not elements — hence the sizeof in every call) are the two frequent stumbling blocks.

  7. A — an assignment between a host and a device array is a synchronous copy with the same synchronisation behaviour as cudaMemcpy.

  8. B — see the two-step data-transfer figure in the pinned-memory episode.

  9. B — cudaMallocHost (or cudaHostAlloc) in C/C++ and the pinned attribute in CUDA Fortran allocate page-locked host memory; cudaPinMemory does not exist.

  10. C — 64 KB per device on all current GPUs; the exact value can be queried through the totalConstMem device property.

  11. C — see the three-step synchronisation figure in the unified-memory episode.