Meeting Summarization Kda

Custom PyTorch Transformer checkpoint trained on MeetingBank for meeting summarization research. This repository is part of the transformer-lab collection.

Model Details

Field Value
Repository Pradheep1647/meeting_summarization_kda-meetingbank-bs8-e20-bf16-4
Attention kda
Dataset meetingbank
Layers 6
Hidden size 512
Heads 8
Batch size 1
Effective batch size 8
Epochs 20
Precision bf16
Checkpoint meeting_model_kda04.pt
Optimizer steps 12,920
Logged training time 59m 46s

Architecture

Architecture diagram

Static architecture diagram generated from this run's config.json, including model width, depth, sequence dimensions, and attention-specific settings.

Training Loss

Training loss

Raw curve data is available in loss_curve.csv.

The curve covers the complete training run. The uploaded checkpoint is the saved epoch with the lowest full-validation loss, not simply the last epoch.

Evaluation

Metric Value
Validation loss 2.5381
Perplexity 12.6559
Token accuracy 0.5497
Top-5 accuracy 0.7400
ROUGE-1 0.2556
ROUGE-2 0.0853
ROUGE-L 0.2055
BLEU 7.90
Evaluation tokens/s 4789.1
Generation tokens/s 92.9
Forward latency (ms) 13.17
Peak CUDA memory (MB) 189.5

Core metrics use the full MeetingBank validation split. Generation metrics use the first 128 validation examples with greedy decoding.

Available Models

Variant Repository
meeting_summarization_kda Pradheep1647/meeting_summarization_kda-meetingbank-bs8-e20-bf16-4

Files

File Purpose
meeting_model_kda04.pt PyTorch checkpoint containing model_state_dict, optimizer states, epoch, and global step.
config.json Training and architecture config converted from the Hydra run config.
architecture.png Architecture diagram generated from the saved model config, with block shapes and dimensions.
tokenizer.json Unified MeetingBank transcript and summary tokenizer.
causal_tokenizer.json Explicit alias of the unified causal tokenizer.
loss_curve.csv TensorBoard train/loss scalar export.
loss_curve.svg Static training-loss plot generated from loss_curve.csv.

Usage

These checkpoints are from a custom PyTorch codebase, not a transformers.AutoModel checkpoint. Use the repo-native builder to instantiate the architecture, then load the checkpoint state dict.

from pathlib import Path

import torch
from huggingface_hub import hf_hub_download
from omegaconf import OmegaConf

import src  # registers components
from src.model.builder import build_causal_lm

repo_id = "Pradheep1647/meeting_summarization_kda-meetingbank-bs8-e20-bf16-4"

config_path = hf_hub_download(repo_id=repo_id, filename="config.json")
checkpoint_path = hf_hub_download(repo_id=repo_id, filename="meeting_model_kda04.pt")

cfg = OmegaConf.load(config_path)
model = build_causal_lm(cfg)

state = torch.load(checkpoint_path, map_location="cpu")
model.load_state_dict(state["model_state_dict"])
model.eval()

print(f"Loaded {repo_id} from {Path(checkpoint_path).name}")

Notes

  • This is a research checkpoint for comparing attention variants under the same MeetingBank setup.
  • The config and tokenizers are included so future runs can reproduce the architecture and preprocessing assumptions.
  • Use config.json as the source of truth for architecture parameters.
Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Pradheep1647/meeting_summarization_kda-meetingbank-bs8-e20-bf16-4

Collection including Pradheep1647/meeting_summarization_kda-meetingbank-bs8-e20-bf16-4

Evaluation results