--- library_name: transformers license: mit base_model: gpt2 tags: - generated_from_trainer - language-modeling - causal-lm - gpt2 - coco model-index: - name: COCO_no_sports_real_severe results: [] language: - en --- # COCO_no_sports_real_severe ## Model Description `COCO_no_sports_real_severe` is a causal language model based on [GPT-2](https://huggingface.co/gpt2), fine-tuned on the florence-generated image captions of a subset of [COCO](https://cocodataset.org/). This subset is labeled for physical activity content in the text: - **Label 0**: Not related to physical activity (e.g., indoor scenes, objects, people at rest) - **Label 1**: Related to physical activity (e.g., sports, exercise, physical activity) The model has been trained on a **general distribution** of this data: - **Label distribution**: `[0.10, 0.90]` This version is designed to serve as the **real model** of our pipeline. Its split corresponds to the **Severe** one. ## Training and Evaluation Data - **Dataset**: [`BeyondDeepfakeDetection/real_train_dataset_v0`](https://huggingface.co/datasets/BeyondDeepfakeDetection/real_train_dataset_v0) - **Label schema**: Binary classification of text as related to physical activity or not. - **Source**: [COCO](https://cocodataset.org/), [Florence]("https://huggingface.co/microsoft/Florence-2-large") ## Training procedure ### Training hyperparameters The following hyperparameters were used during training: - learning_rate: 2e-05 - train_batch_size: 8 - eval_batch_size: 16 - seed: 42 - optimizer: Use adamw_torch with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments - lr_scheduler_type: linear - lr_scheduler_warmup_steps: 1000 - num_epochs: 5 - mixed_precision_training: Native AMP ### Training results | Training Loss | Epoch | Step | Validation Loss | |:-------------:|:-----:|:----:|:---------------:| | 1.2429 | 1.0 | 1196 | 0.9963 | | 0.9985 | 2.0 | 2392 | 0.8874 | | 0.901 | 3.0 | 3588 | 0.8470 | | 0.87 | 4.0 | 4784 | 0.8288 | | 0.837 | 5.0 | 5980 | 0.8214 | ### Framework versions - Transformers 4.46.3 - Pytorch 2.1.2+cu121 - Datasets 2.19.1 - Tokenizers 0.20.3 ## Get started In order to infer the joint probability of phrases under this model you can use the following code: ```python from transformers import AutoTokenizer, AutoModelForCausalLM import torch import torch.nn.functional as F import pandas as pd from huggingface_hub import login from tqdm import tqdm from datasets import load_dataset # Define variables hf_token = "" model_name = f"BeyondDeepFakeDetection/COCO_no_sports_real_v3" text_column = "text" dataset = "BeyondDeepFakeDetection/COCO_no_sports" # Load Model tokenizer = AutoTokenizer.from_pretrained("gpt2") model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto") device = "cuda" if torch.cuda.is_available() else "cpu" tokenizer.pad_token = tokenizer.eos_token model.to(device) # Login login(token=hf_token) def compute_log_probabilities_for_sequence(model, tokenizer, input_text): inputs = tokenizer(input_text, return_tensors="pt", padding=True, truncation=True).to(device) input_ids = inputs["input_ids"] attention_mask = inputs["attention_mask"] with torch.no_grad(): outputs = model(input_ids=input_ids, attention_mask=attention_mask) logits = outputs.logits[:, :-1, :] target_ids = input_ids[:, 1:] log_probs = F.log_softmax(logits, dim=-1) seq_token_logprobs = log_probs.gather(2, target_ids.unsqueeze(-1)).squeeze(-1) word_probabilities = [] for i, token_id in enumerate(target_ids[0]): word = tokenizer.decode([token_id]) log_prob = seq_token_logprobs[0, i].item() word_probabilities.append((word, log_prob)) return word_probabilities test_df = pd.DataFrame(load_dataset(dataset, split="train")) results = [] for count, text in enumerate(tqdm(test_df[text_column], desc="Processing Texts")): word_probs = compute_log_probabilities_for_sequence(model, tokenizer, text) total_log_prob = sum(prob for _, prob in word_probs) avg_log_prob = total_log_prob / len(word_probs) if word_probs else float("-inf") results.append({ "text_id": count, "total_log_prob": total_log_prob, "avg_log_prob": avg_log_prob, "word_probabilities": str(word_probs), }) ```