Show HN: TurboGPT: train 22KiB transformer in 13s A developer released TurboGPT, an MIT-licensed byte-level GPT trainer written in CUDA C++ that trains a 22KiB transformer in 13 seconds. The tool builds on Windows with Visual Studio 2022 C++ tools and CUDA 13.4 via `build.ps1 -CudaArch 86`, stores checkpoints with model, optimizer, scheduler and trainer state for resuming with `--load`, and emits TensorBoard-compatible logs plus a `report.json`; its author reports 2.52435 BPB on hn1g after 1.5G training tokens. Tiny byte-level GPT training in CUDA C++. MIT. Windows, Visual Studio 2022 C++ tools, and CUDA 13.4: .\build.ps1 -CudaArch 86 CudaArch is the GPU compute capability from NVIDIA's CUDA GPU list https://developer.nvidia.com/cuda-gpus . .\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4 The run stores its checkpoint at runs/ctx4/ctx4.pt , containing model, optimizer, scheduler, and trainer state. Use --load CHECKPOINT.pt to resume it. runs/ctx4/report.json is derived from the log directory. Logs are TensorBoard-compatible: one report per batch, capped at 8Mi reports, and flushed with periodic or final checkpoints. - hn1g after 1.5G training tokens: 2.52435 BPB . python tests\verify.py