CLIP-ViT-B-16-EntityNet-33M-Finetuned

A CLIP ViT-B/16 model, initialized from the DataComp-1B CLIP (ViT-B-16/datacomp_xl_s13b_b90k) and finetuned on EntityNet-33M.

Finetuning on EntityNet keeps the base model's ImageNet accuracy and more than doubles its accuracy on fine-grained species recognition, with a small drop in image-text retrieval and on ImageNet distribution shifts.

Paper · Code · Dataset · Project page

Results

Zero-shot top-1 accuracy (%), all models evaluated with the same protocol from the paper, so numbers can differ slightly from those on other model cards. Retrieval is the average image-text recall@1 over COCO, Flickr30k and XM3600. Rare Species tests 400 species that were removed from our training data. Following the paper, it is only reported for models whose training data is known to exclude them ("–" otherwise).

Model ImageNet iNat 2021 CUB Rare Species Retrieval
Base model: DataComp-1B ViT-B/16 73.5 15.3 79.0 – 57.4
This model 73.5 34.9 86.5 – 52.2

Usage

import torch
import open_clip
from PIL import Image

model, _, preprocess = open_clip.create_model_and_transforms("hf-hub:lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned")
model.eval()
tokenizer = open_clip.get_tokenizer("ViT-B-16", context_length=32)  # trained with 32 text tokens

image = preprocess(Image.open("image.jpg")).unsqueeze(0)
labels = ["a dog", "a cat", "a rabbit"]
text = tokenizer(labels)

with torch.no_grad():
    image_features = model.encode_image(image)
    text_features = model.encode_text(text)
    image_features /= image_features.norm(dim=-1, keepdim=True)
    text_features /= text_features.norm(dim=-1, keepdim=True)
    probs = (100.0 * image_features @ text_features.T).softmax(dim=-1)

print(labels[probs.argmax().item()])

Species names work in English ("red fox") and Latin ("Vulpes vulpes"); the paper reports the better of the two.

Training

The model has seen ~13B images at a batch size of 90k during pretraining, and ~0.6B images at a batch size of 32k during finetuning. Text labels were sampled 50/50 from the images' web alt texts and from the knowledge-graph text behind the image (the search query, or the entity's aliases or descriptions). See the paper for details.

All EntityNet models

Model Training Best for
CLIP-ViT-B-16-EntityNet-33M from scratch, all 33M images studying CLIP training with full control over the data
CLIP-ViT-B-32-EntityNet-33M from scratch, all 33M images same, smaller and faster
CLIP-ViT-B-16-EntityNet-33M-Finetuned DataComp-1B finetuned on 33M general use with a boost on animals and plants
CLIP-ViT-B-32-EntityNet-33M-Finetuned DataComp-1B finetuned on 33M same, smaller and faster
CLIP-ViT-B-16-LivingThings-10M-Finetuned DataComp-1B finetuned on 10M organisms best species recognition
CLIP-ViT-B-32-LivingThings-10M-Finetuned DataComp-1B finetuned on 10M organisms species recognition, smaller and faster
CLIP-ViT-B-32-LivingThings-10M from scratch, 10M organisms a species expert trained from scratch on organism images only

Citation

@inproceedings{ging2026entitynet,
  author    = {Simon Ging and Sebastian Walter and Jelena Bratuli{\'c} and Johannes Dienert and Hannah Bast and Thomas Brox},
  title     = {Using Knowledge Graphs to Harvest Datasets for Efficient {CLIP} Model Training},
  booktitle = {Pattern Recognition, 47th {DAGM} German Conference, {DAGM} {GCPR} 2025, Freiburg, Germany, September 23--26, 2025, Proceedings},
  series    = {Lecture Notes in Computer Science},
  publisher = {Springer Nature Switzerland},
  year      = {2026},
  pages     = {287--302},
  isbn      = {978-3-032-12840-9},
  doi       = {10.1007/978-3-032-12840-9_19}
}
Downloads last month
51
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned

Finetuned
(2)
this model

Dataset used to train lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned

Collection including lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned

Paper for lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned