ANLP Assignment 2

Links

Setup

Dependencies are pinned in pyproject.toml / uv.lock:

uv sync

Tasks 1 and 3 load their datasets from Hugging Face. For Task 2, save browndw/human-ai-parallel-corpus as a Hugging Face Dataset with a text column at src/part2/dataset/.

Usage

Run from the repository root:

uv run python -m src.part1.train --all          # train the five FFN variants
uv run python -m src.part1.evaluate             # evaluate checkpoints and plot expert usage
uv run python -m src.part2.train --all          # train with all five optimizers
uv run python -m src.part3.evaluate --num_samples 1000 --max_new_tokens 40

Task 1 checkpoints are written to checkpoints_part1/best_variant_<n>.pt, and evaluation results to assets/part1_results/. Task 2 logs metrics to Weights & Biases and saves best.pt, last.pt, and metrics.json under checkpoints_part2/<optimizer>/<run_id>/. The weights are PyTorch state dictionaries; metrics.json includes the model configuration for loading them. Task 3 results are written to task3_results.json.

Task 3 uses the original pretrained EleutherAI/pythia-160m checkpoint.

Loading a Task 2 checkpoint

These are state dictionaries for the custom PyTorch architecture in this repository. Clone or download the source files, install the dependencies, and run from that checkout:

import json
import torch
from huggingface_hub import hf_hub_download
from src.part2.model import DecoderOnlyTransformer

repo_id = "Chechaaa09/anlp-a2-models"
folder = "checkpoints_part2/adamw/oj3opt7s"
with open(hf_hub_download(repo_id, f"{folder}/metrics.json")) as handle:
    metadata = json.load(handle)
model = DecoderOnlyTransformer(**metadata["model_config"])
weights = torch.load(
    hf_hub_download(repo_id, f"{folder}/best.pt"),
    map_location="cpu", weights_only=True,
)
model.load_state_dict(weights)
model.eval()

Use the GPT-2 tokenizer for Task 2. Each run includes best.pt (lowest validation loss), last.pt (final weights), and metrics.json (configuration, all ten evaluations, and its W&B link). Task 1 uses the saved custom tokenizer in tokenizer_part1/.

Optimizer Checkpoint folder Validation loss Validation perplexity Test sentence BLEU
adamw checkpoints_part2/adamw/oj3opt7s/ 5.2531 191.15 2.9645
cautious checkpoints_part2/cautious/p2yw4r11/ 5.4965 243.84 2.5585
lion checkpoints_part2/lion/ydyarkxk/ 4.9409 139.90 2.5707
scion checkpoints_part2/scion/89yiklpr/ 5.5330 252.91 2.7522
sophia checkpoints_part2/sophia/3j32zwf3/ 5.3892 219.02 2.5534

Task 2 results are from the completed checkpoint-saving runs on 4 October 2026. BLEU is on the 0–100 scale.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support