Assignment 2 Part 2: ADAMW
Custom dense decoder-only PyTorch model (27,303,936 unique parameters), trained from fresh initialization for one complete pass with the ADAMW optimizer. The checkpoint records 1,219 updates and 48,559,691 supervised tokens. It contains model weights, architecture configuration, optimizer name, training progress, and data identity; it does not contain optimizer state.
The training source is browndw/human-ai-parallel-corpus. The shared frozen
16,000-token byte-level BPE tokenizer and self-contained model source are bundled.
Requires PyTorch 2.7.0 and tokenizers 0.23.1.
Download this repository, add its directory to sys.path, then call
from load_model import load_model; model, tokenizer = load_model(directory).
model.forward(input_ids, attention_mask) returns logits in a dictionary.
Final validation CE: 4.072847 nats per supervised token.
Final test corpus BLEU: 0.919434 on 3,319 fixed
64-token continuations. SacreBLEU signature:
nrefs:1|case:mixed|eff:no|tok:13a|smooth:exp|version:2.5.1.
- Downloads last month
- 226