cd/sources/gilesthomas-auto-discovered· home› sources› Gilesthomas (auto-discovered)
cat /sources/gilesthomas-auto-discovered.feed | wc -l → 23

Gilesthomas (auto-discovered)

articles 23 domain gilesthomas.com → page 1/2 feed RSS
16:40
2026-10-01
gilesthomas.com
large-language-models

Why do OpenAI's GPT-2 weights beat mine? Part five: data quality

Independent LLM-from-scratch experiments found that OpenAI's GPT-2 small (124M parameters) consistently outperformed the author's 163M-parameter models on an instruction fine-tuning task adapted from …

17:25
2026-09-03
gilesthomas.com
artificial-intelligence

Putting my JAX-trained models on the Hugging Face Hub

Developer gpjt has uploaded PyTorch-compatible versions of all his JAX-trained GPT-2 models to the Hugging Face Hub, including models from his blog series on writing an LLM from scratch and Chinchilla…

22:00
2026-08-25
gilesthomas.com
developer-tools

Adding diagrams to my static site generator with D2

Simon Willison added D2 diagram support to his static site generator, allowing him to compile D2 files into SVG diagrams for blog posts. The feature, implemented via a Python function that invokes the…

01:12
2026-08-20
gilesthomas.com
machine-learning

Use the built-in GELU, don't roll your own!

PyTorch's built-in GELU function is 20% faster than a hand-rolled version when training GPT-2 small models, according to a developer's benchmark. The same code training the same model on the same data…

19:00
2026-08-14
gilesthomas.com
artificial-intelligence

Why do OpenAI's GPT-2 weights beat mine? Part two: IFT dropout

OpenAI's GPT-2 medium weights achieved the highest IFT score of 42.43 with only 2 IFT epochs, while a JAX model with no dropout and no MHA bias scored 21.45 with 5 epochs, and a JAX model with dropout…

19:00
2026-08-07
gilesthomas.com
machine-learning

A quick(ish) Chinchilla check

Giles Thomas, a developer, tested the Chinchilla scaling rule by comparing overtrained GPT-2 style models (trained on 40 tokens per parameter) against a model scaled up in parameters and tokens equall…

19:31
2026-07-31
gilesthomas.com
artificial-intelligence

I use AI on this blog

In a blog post, the author describes their personal policy for using AI tools like ChatGPT and Claude in creating content for their blog, emphasizing that AI is used for ideation, code review, and edi…

18:13
2026-07-30
gilesthomas.com
large-language-models

Why do OpenAI's GPT-2 weights beat mine? Part two: the bugfix

A bug in the evaluation code for GPT-2 style models caused incorrect baseline numbers, but OpenAI's original weights still outperform the author's models on instruction-following tasks. The bug involv…

15:00
2026-07-29
gilesthomas.com
large-language-models

Why do OpenAI's GPT-2 weights beat mine?

OpenAI's original GPT-2 small and medium weights consistently outperform custom-trained models in instruction-following evaluations, despite some custom models achieving better test loss, according to…

22:33
2026-07-24
gilesthomas.com
large-language-models

Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090

Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 using Unsloth's UD-IQ4_NL_XL quantisation achieved up to 140 tokens per second for generation and over 3,300 tok/s for prompt processing with a…

20:49
2026-07-10
gilesthomas.com
large-language-models

Building intuition about LLM parameter counts

A developer building a GPT-2 implementation in JAX discovered that token embeddings and the output head account for nearly half of the model's 163 million parameters, while attention layers use fewer …

00:23
2026-07-09
gilesthomas.com
artificial-intelligence

poppy the training box, part 1: the beginnings

A developer repurposed an old small-form-factor PC named 'poppy' into a dedicated machine for local LLM training, upgrading its case and power supply to accommodate future multi-GPU setups. The projec…

20:15
2026-06-24
gilesthomas.com
large-language-models

Thoughts on Role Confusion

Researchers Charles Ye, Jasmine Cui, and Dylan Hadfield-Menell found that large language models often ignore explicit role tags like <system> or <user> and instead infer roles from text tone, enabling…

02:11
2026-06-17
gilesthomas.com
machine-learning

Flax debugging: making a hash of things

A developer debugging a JAX/Flax NNX training loop discovered that the loss was stuck at 10.82, indicating the model was performing no better than random guessing. The issue was traced to the training…

20:40
2026-06-15
gilesthomas.com
machine-learning

Jax: Commitment Issues

JAX's default_device context manager places arrays on the specified device but does not commit them, allowing JAX to move them to other devices. This caused array lookups to take over a second by trig…

page 1 / 2 next →