Vito Lin

First CUDA program: vector addition

· 1 minute read · Discussion on LinkedIn

Today I implemented a CUDA program — vector addition on GPU.

Host vs. device

Host = CPU. Device = GPU.

Memory flow matters

Unlike typical CPU programs, CUDA requires explicit memory management:

  1. Allocate memory on CPU
  2. Allocate memory on GPU
  3. Copy data to GPU
  4. Execute kernel
  5. Copy results back
  6. Free memory

Memory is everything in CUDA

Performance in CUDA is heavily tied to memory allocation, CPU ↔ GPU transfer, and memory access patterns.

Source code: cuda_add.cu

Vector addition running on the GPU