Instructions to use lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- OpenCLIP
How to use lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned with OpenCLIP:
import open_clip model, preprocess_train, preprocess_val = open_clip.create_model_and_transforms('hf-hub:lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned') tokenizer = open_clip.get_tokenizer('hf-hub:lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned') - Notebooks
- Google Colab
- Kaggle
CLIP-ViT-B-16-EntityNet-33M-Finetuned
A CLIP ViT-B/16 model, initialized from the DataComp-1B CLIP (ViT-B-16/datacomp_xl_s13b_b90k) and finetuned on EntityNet-33M.
Finetuning on EntityNet keeps the base model's ImageNet accuracy and more than doubles its accuracy on fine-grained species recognition, with a small drop in image-text retrieval and on ImageNet distribution shifts.
Paper · Code · Dataset · Project page
Results
Zero-shot top-1 accuracy (%), all models evaluated with the same protocol from the paper, so numbers can differ slightly from those on other model cards. Retrieval is the average image-text recall@1 over COCO, Flickr30k and XM3600. Rare Species tests 400 species that were removed from our training data. Following the paper, it is only reported for models whose training data is known to exclude them ("–" otherwise).
| Model | ImageNet | iNat 2021 | CUB | Rare Species | Retrieval |
|---|---|---|---|---|---|
| Base model: DataComp-1B ViT-B/16 | 73.5 | 15.3 | 79.0 | – | 57.4 |
| This model | 73.5 | 34.9 | 86.5 | – | 52.2 |
Usage
import torch
import open_clip
from PIL import Image
model, _, preprocess = open_clip.create_model_and_transforms("hf-hub:lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned")
model.eval()
tokenizer = open_clip.get_tokenizer("ViT-B-16", context_length=32) # trained with 32 text tokens
image = preprocess(Image.open("image.jpg")).unsqueeze(0)
labels = ["a dog", "a cat", "a rabbit"]
text = tokenizer(labels)
with torch.no_grad():
image_features = model.encode_image(image)
text_features = model.encode_text(text)
image_features /= image_features.norm(dim=-1, keepdim=True)
text_features /= text_features.norm(dim=-1, keepdim=True)
probs = (100.0 * image_features @ text_features.T).softmax(dim=-1)
print(labels[probs.argmax().item()])
Species names work in English ("red fox") and Latin ("Vulpes vulpes"); the paper reports the better of the two.
Training
The model has seen ~13B images at a batch size of 90k during pretraining, and ~0.6B images at a batch size of 32k during finetuning. Text labels were sampled 50/50 from the images' web alt texts and from the knowledge-graph text behind the image (the search query, or the entity's aliases or descriptions). See the paper for details.
All EntityNet models
| Model | Training | Best for |
|---|---|---|
| CLIP-ViT-B-16-EntityNet-33M | from scratch, all 33M images | studying CLIP training with full control over the data |
| CLIP-ViT-B-32-EntityNet-33M | from scratch, all 33M images | same, smaller and faster |
| CLIP-ViT-B-16-EntityNet-33M-Finetuned | DataComp-1B finetuned on 33M | general use with a boost on animals and plants |
| CLIP-ViT-B-32-EntityNet-33M-Finetuned | DataComp-1B finetuned on 33M | same, smaller and faster |
| CLIP-ViT-B-16-LivingThings-10M-Finetuned | DataComp-1B finetuned on 10M organisms | best species recognition |
| CLIP-ViT-B-32-LivingThings-10M-Finetuned | DataComp-1B finetuned on 10M organisms | species recognition, smaller and faster |
| CLIP-ViT-B-32-LivingThings-10M | from scratch, 10M organisms | a species expert trained from scratch on organism images only |
Citation
@inproceedings{ging2026entitynet,
author = {Simon Ging and Sebastian Walter and Jelena Bratuli{\'c} and Johannes Dienert and Hannah Bast and Thomas Brox},
title = {Using Knowledge Graphs to Harvest Datasets for Efficient {CLIP} Model Training},
booktitle = {Pattern Recognition, 47th {DAGM} German Conference, {DAGM} {GCPR} 2025, Freiburg, Germany, September 23--26, 2025, Proceedings},
series = {Lecture Notes in Computer Science},
publisher = {Springer Nature Switzerland},
year = {2026},
pages = {287--302},
isbn = {978-3-032-12840-9},
doi = {10.1007/978-3-032-12840-9_19}
}
- Downloads last month
- 51
Model tree for lmb-freiburg/CLIP-ViT-B-16-EntityNet-33M-Finetuned
Base model
laion/CLIP-ViT-B-16-DataComp.XL-s13B-b90K