ANLP Assignment 2
Links
- Weights & Biases: Task 1, Task 2, Task 3.
- Checkpoints (Hugging Face): https://huggingface.co/Chechaaa09/anlp-a2-models
Setup
Dependencies are pinned in pyproject.toml / uv.lock:
uv sync
Tasks 1 and 3 load their datasets from Hugging Face. For Task 2, save
browndw/human-ai-parallel-corpus as a Hugging Face Dataset with a text
column at src/part2/dataset/.
Usage
Run from the repository root:
uv run python -m src.part1.train --all # train the five FFN variants
uv run python -m src.part1.evaluate # evaluate checkpoints and plot expert usage
uv run python -m src.part2.train --all # train with all five optimizers
uv run python -m src.part3.evaluate --num_samples 1000 --max_new_tokens 40
Task 1 checkpoints are written to checkpoints_part1/best_variant_<n>.pt,
and evaluation results to assets/part1_results/. Task 2 logs metrics to
Weights & Biases and saves best.pt, last.pt, and metrics.json under
checkpoints_part2/<optimizer>/<run_id>/. The weights are PyTorch state
dictionaries; metrics.json includes the model configuration for loading them.
Task 3 results are written to task3_results.json.
Task 3 uses the original pretrained EleutherAI/pythia-160m checkpoint.
Loading a Task 2 checkpoint
These are state dictionaries for the custom PyTorch architecture in this repository. Clone or download the source files, install the dependencies, and run from that checkout:
import json
import torch
from huggingface_hub import hf_hub_download
from src.part2.model import DecoderOnlyTransformer
repo_id = "Chechaaa09/anlp-a2-models"
folder = "checkpoints_part2/adamw/oj3opt7s"
with open(hf_hub_download(repo_id, f"{folder}/metrics.json")) as handle:
metadata = json.load(handle)
model = DecoderOnlyTransformer(**metadata["model_config"])
weights = torch.load(
hf_hub_download(repo_id, f"{folder}/best.pt"),
map_location="cpu", weights_only=True,
)
model.load_state_dict(weights)
model.eval()
Use the GPT-2 tokenizer for Task 2. Each run includes best.pt (lowest
validation loss), last.pt (final weights), and metrics.json (configuration,
all ten evaluations, and its W&B link). Task 1 uses the saved custom tokenizer
in tokenizer_part1/.
| Optimizer | Checkpoint folder | Validation loss | Validation perplexity | Test sentence BLEU |
|---|---|---|---|---|
| adamw | checkpoints_part2/adamw/oj3opt7s/ |
5.2531 | 191.15 | 2.9645 |
| cautious | checkpoints_part2/cautious/p2yw4r11/ |
5.4965 | 243.84 | 2.5585 |
| lion | checkpoints_part2/lion/ydyarkxk/ |
4.9409 | 139.90 | 2.5707 |
| scion | checkpoints_part2/scion/89yiklpr/ |
5.5330 | 252.91 | 2.7522 |
| sophia | checkpoints_part2/sophia/3j32zwf3/ |
5.3892 | 219.02 | 2.5534 |
Task 2 results are from the completed checkpoint-saving runs on 4 October 2026. BLEU is on the 0–100 scale.