(Advanced) CUDA Programming Course

Welcome to the CUDA Programming Course. This course covers high-performance kernel development for modern NVIDIA GPUs, from core concepts to the newer features in Ampere, Hopper, and Blackwell architectures.

The exercises are downloaded and run against your own GPU: how the exercises work. No GPU of your own? The gpu-submit toolkit runs your CUDA on a datacentre GPU from your laptop.

Course Outline

Part 1 — Introduction

Part 2 — Basic Kernels

Part 3 — Divergence and Coalescing

Part 4 — Pipelining and Occupancy

Part 5 — Shared Memory

Part 6 — Thread Coarsening and Vectorized Memory Access

Part 7 — Warp Shuffles, Reductions, and Cooperative Groups

Part 8 — Asynchronous Data Movement: LDGSTS

Glossary.