Quiz solutions¶
Answers
B — the “small, read-only, broadcast to many threads” pattern is what constant memory is for.
A —
cudaMallocManaged/managedneeds no explicit copy. B works but needs two explicit copies. C cannot work: host code cannot dereference acudaMallocpointer, so there is no “host side” to do arithmetic on. D leaves the two buffers unrelated — the device buffer is never filled and the result never comes back (and plainmallocmemory is not accessible to kernels without unified memory).B — automatic page migration is the defining feature of managed memory.
C — “managed” does not imply “synchronised”.
A — the canonical workflow of explicit (manual) memory management.
A — the argument order (destination first) and the unit of
count(bytes, not elements — hence thesizeofin every call) are the two frequent stumbling blocks.A — an assignment between a host and a
devicearray is a synchronous copy with the same synchronisation behaviour ascudaMemcpy.B — see the two-step data-transfer figure in the pinned-memory episode.
B —
cudaMallocHost(orcudaHostAlloc) in C/C++ and thepinnedattribute in CUDA Fortran allocate page-locked host memory;cudaPinMemorydoes not exist.C — 64 KB per device on all current GPUs; the exact value can be queried through the
totalConstMemdevice property.C — see the three-step synchronisation figure in the unified-memory episode.