Skip to content

Repository files navigation

1. Project structure

GPT-2/
├── gpt2/
│   ├── __init__.py
│   ├── utils.py          # seeds, devices, token windows, loaders, checkpoints
│   ├── model.py          # GPT-2 config, attention, blocks, and language model
│   ├── training.py       # reusable Trainer and gradient accumulation
│   ├── evaluation.py     # loss, perplexity, greedy/top-k generation
│   └── weights.py        # OpenAI checkpoint download and TensorFlow mapping
├── scripts/
│   ├── train.py          # train from a local text corpus
│   ├── evaluate.py       # evaluate/generate from a local PyTorch checkpoint
│   └── generate.py       # generate with original OpenAI GPT-2 weights
├── pyproject.toml
├── requirements.txt
└── requirements-dev.txt

2. What each module does

gpt2/utils.py

Provides a tokenizer protocol, reproducible seeding, automatic device selection, text/token conversion, a sliding next-token dataset, DataLoader construction, parameter counting, and atomic checkpoint save/load functions.

gpt2/model.py

Implements GPT-2 without external chapter files:

  • token and positional embeddings;
  • fused query/key/value causal multi-head attention;
  • pre-normalized transformer blocks;
  • GPT-2 GELU and feed-forward network;
  • tied token/output embeddings;
  • optional causal-language-model loss.

gpt2/training.py

Contains TrainConfig, TrainingHistory, and Trainer. Gradient accumulation is token-weighted, so an incomplete final accumulation group is normalized correctly. It also supports gradient clipping, schedulers, periodic evaluation, and best/final checkpoints.

gpt2/evaluation.py

Contains token-weighted evaluation loss, perplexity, and batch-safe generation. Generation supports greedy decoding, temperature, top-k filtering, context cropping, and per-sequence EOS handling.

gpt2/weights.py

Downloads original OpenAI GPT-2 checkpoint files, imports TensorFlow lazily, reads the checkpoint arrays, and maps them into the fused PyTorch architecture. Weights are copied into existing parameters, preserving optimizer references, device placement, and dtype.

3. Environment setup

Python 3.11 is recommended.

conda create -n modular-gpt2 python=3.11 -y
conda activate modular-gpt2
python -m pip install --upgrade pip

Install the PyTorch build appropriate for your machine first. For CPU-only:

python -m pip install --no-cache-dir torch \
  --index-url https://download.pytorch.org/whl/cpu

Then, from the project directory:

python -m pip install --no-cache-dir -e ".[dev,tokenizer]"

Only when you need to import the original OpenAI TensorFlow checkpoints:

python -m pip install --no-cache-dir -e ".[openai-weights]"

TensorFlow and the original checkpoint files are large. Leave this optional component uninstalled when training from scratch or using only local PyTorch checkpoints.

4. Train from scratch

Use any sufficiently long UTF-8 text file:

python scripts/train.py path/to/corpus.txt \
  --device auto \
  --epochs 2 \
  --batch-size 4 \
  --context-length 128

The script saves:

checkpoints/best.pt
checkpoints/final.pt

For a quick CPU smoke test, use a smaller model:

python scripts/train.py path/to/corpus.txt \
  --device cpu \
  --epochs 1 \
  --batch-size 2 \
  --context-length 32 \
  --emb-dim 64 \
  --heads 4 \
  --layers 2 \
  --eval-every 2

5. Evaluate a local checkpoint

Evaluate loss/perplexity and generate a sample:

python scripts/evaluate.py checkpoints/final.pt \
  --text-file path/to/evaluation.txt \
  --prompt "Every effort moves you" \
  --device auto

Generate only by omitting --text-file:

python scripts/evaluate.py checkpoints/final.pt \
  --prompt "Every effort moves you"

6. Load original OpenAI GPT-2 weights

After installing the openai-weights optional dependency:

python scripts/generate.py \
  --model-size 124M \
  --prompt "Every effort moves you" \
  --device auto

The first run downloads the original files into gpt2_models/124M. Later runs skip files whose local size matches the server's content length.

Attribution

This refactor is based on the Apache-2.0-licensed educational code accompanying Sebastian Raschka's Build a Large Language Model From Scratch.

About

A GPT-2 implementation in PyTorch with training, evaluation, generation, and pretrained-weight loading.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages