Writing
- 2026 · 08Pretraining and fine-tuning nanoGPT-124M from scratch
- 2026 · 08Paper notes: The Llama 3 Herd of Models
- 2026 · 07Paper notes: PagedAttention
- 2026 · 07Paper notes: Trading More Storage for Less Computation
- 2026 · 07Resources that helped me transition into AI Infrastructure
- 2026 · 05Register tiling: the foundation of high-performance CUDA GEMM
- 2026 · 05Becoming a contributor to NVIDIA's cuda-oxide
- 2026 · 05How CUDA is executed by the GPU
- 2026 · 04Why adding just 1 block can drop GPU utilization to 67%
- 2026 · 04What GPU hardware really taught me
- 2026 · 04How to practice CUDA without a GPU (for free)
- 2026 · 04CUDA matrix multiplication: optimizing with shared memory and tiling
- 2026 · 03CUDA's memory model: why memory, not compute, is the bottleneck
- 2026 · 03CUDA's execution model: threads, warps, and how the GPU schedules them
- 2026 · 03First CUDA program: vector addition