SIMT¶
SIMT means Single Instruction, Multiple Threads.
In CUDA, threads are grouped into warps, usually 32 threads. Threads in a warp execute the same instruction at the same time, but each thread has its own: - registers - thread ID - memory addresses - program state Example: vector addition.
Bank¶
In CUDA, a memory bank is an independent subdivision of shared memory that can serve an access in parallel with other banks.
Modern NVIDIA GPUs typically have 32 shared-memory banks, matching the 32 threads in a warp.
A bank conflict occurs when multiple threads access different addresses in the same bank:
Those accesses may be serialized, reducing performance.Exception: if multiple threads read the same address, the hardware can usually broadcast the value without a conflict.