- 2026 · 08Pretraining and fine-tuning nanoGPT-124M from scratch
- 2026 · 08Paper notes: The Llama 3 Herd of Models
- 2026 · 07Paper notes: PagedAttention
- 2026 · 07Paper notes: Trading More Storage for Less Computation
- 2026 · 07Resources that helped me transition into AI Infrastructure
- 2026 · 05Register tiling: the foundation of high-performance CUDA GEMM
- 2026 · 05Becoming a contributor to NVIDIA's cuda-oxide
- 2026 · 05How CUDA is executed by the GPU
- 2026 · 04Why adding just 1 block can drop GPU utilization to 67%
- 2026 · 04What GPU hardware really taught me
- 2026 · 04How to practice CUDA without a GPU (for free)
- 2026 · 04CUDA matrix multiplication: optimizing with shared memory and tiling
- 2026 · 03CUDA's memory model: why memory, not compute, is the bottleneck
- 2026 · 03CUDA's execution model: threads, warps, and how the GPU schedules them
- 2026 · 03First CUDA program: vector addition
- cs-notes—A collection of computer science notes.·6 ★
- cuda_practice—Some code about CUDA·5 ★
- vito—A bookkeeping system written in Rust·3 ★
- EasySwitch—
- nanoGPT_124m_practice—Inspired from Andrej karpathy's nanogpt repo
Pull Requests
- NVIDIA/cuda-rust—5 patches
- NVlabs/cutile-rs—2 patches
- google/rrg—1 patch
- lgadi/rust_proc_list—1 patch
- TheExplainthis/ChatGPT-Line-Bot—1 patch