cd /news/large-language-models/building-a-transformer-from-scratch-… · home › topics › large-language-models › article
[ARTICLE · art-148713] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Building a Transformer from Scratch: Part 1 — The Embedding Layer

A developer published a from-scratch PyTorch implementation of the embedding layer used in autoregressive Transformers like GPT-2, combining a learned token lookup table with learned positional embeddings via element-wise addition. The writeup explains why raw integer token IDs must be projected into continuous vectors, why nn.Embedding is a direct row lookup equivalent to one-hot matrix multiplication, and includes a guard clause that raises an error when sequence length exceeds the maximum context window.

by read3 min views2 publishedOct 10, 2026

In an autoregressive Transformer such as GPT, the work begins before the attention mechanism processes a single tensor. A model cannot operate on raw text, and integer token IDs carry no semantic structure on their own. The embedding layer bridges this gap by converting discrete token IDs into continuous, high-dimensional vectors that encode both the meaning of each token and its position in the sequence.

This article explains how the embedding layer works and presents a clean PyTorch implementation.

The Problem Embeddings Solve

A tokenizer splits raw text into segments and maps each one to an integer index. For example, the word "transformer" might map to 41551 (an illustrative value).

Passing these integers directly into a neural network is problematic, because numerical values imply an ordinal relationship that does not exist. Token 41552 is not "greater than" token 41551 in any meaningful sense.

The solution uses two learned lookup tables:

W_e) d_model. Through backpropagation, tokens that appear in similar linguistic contexts converge toward similar regions of the vector space.W_p) 0, 1, ..., T-1 is therefore mapped to a learned vector of dimension The module below combines token embeddings with learned positional embeddings, following the architecture used in GPT-2.

import torch
import torch.nn as nn

class TransformerEmbeddings(nn.Module):
    def __init__(self, vocab_size: int, d_model: int, max_seq_len: int):
        super().__init__()
        self.token_embeddings = nn.Embedding(vocab_size, d_model)

        self.position_embeddings = nn.Embedding(max_seq_len, d_model)

    def forward(self, input_ids: torch.Tensor) -> torch.Tensor:
        batch_size, seq_len = input_ids.shape

        if seq_len > self.position_embeddings.num_embeddings:
            raise ValueError(
                f"Sequence length ({seq_len}) exceeds maximum context window "
                f"({self.position_embeddings.num_embeddings})"
            )

        positions = torch.arange(0, seq_len, device=input_ids.device).unsqueeze(0)

        tok_emb = self.token_embeddings(input_ids)      # (batch_size, seq_len, d_model)
        pos_emb = self.position_embeddings(positions)   # (1, seq_len, d_model)

        return tok_emb + pos_emb

nn.Embedding Works nn.Embedding is a trainable weight matrix:

W ∈ R^(vocab_size × d_model)

It is mathematically equivalent to multiplying a one-hot vector by this matrix (x_one_hot @ W), but it is implemented as a direct row lookup. This avoids constructing large, sparse one-hot tensors and removes the associated memory and compute overhead.

device=input_ids.device ensures that the position indices are created on the same device (CPU, CUDA, or Apple Silicon MPS) as the input. Omitting it defaults to the CPU and raises a device mismatch error when training on a GPU.(1, seq_len, d_model), while tok_emb has shape (batch_size, seq_len, d_model). When the two are added, PyTorch broadcasts the position vectors across every sequence in the batch. Concatenating the two vectors would widen each token representation to 2 × d_model, increasing the parameter count and memory usage of every subsequent linear projection.

Element-wise addition preserves the channel dimension at d_model. Conceptually, the positional embedding acts as an offset applied to the token vector, allowing the model to distinguish the same token at different positions.

The combined embeddings leave this layer with shape (batch_size, seq_len, d_model) and feed directly into the first Transformer block. The next article in this series covers the attention mechanism.

── more in #large-language-models 4 stories · sorted by recency
── more on @pytorch 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/building-a-transform…] indexed:0 read:3min 2026-10-10 · —