What a Shot: film shot-size classifiers
Small vision transformers that read a film frame's shot size, from extreme long shot to detail, exported to ONNX. They run on CPU or in the browser (ONNX Runtime Web); no PyTorch needed.
- Try them: What a Shot tags movie trailers live in the browser with v6, and shows every version's prediction and attention on human-labelled test frames.
- See them used: AiRoll cuts clips to the beat of a song and labels each
shot with
dinov3_shot_v5_int8.onnx. - Data: types-of-film-shots: 53,449 labelled training rows (48,309 distinct stills from about 1,000 films) and 863 human-labelled test frames.
- License: CC BY-NC 4.0: non-commercial use only, and credit is required, also for fine-tuned models.
Live on movie trailers
In the What a Shot Space v6 runs on every frame it can while a trailer plays and tags the shot size in the corner; the timeline below draws its answers over v5's. v6 is, a MobileNetV3-L distilled from v5, running in real time in the browser: about 50 ms a frame (about 20 frames a second) on a desktop CPU, in WebAssembly on one thread.
Disclaimer: the trailers are shown for educational and research purposes only, to demonstrate the classifier. I do not own them: all rights belong to their studios and publishers. If you hold the rights to one and want it removed, open a discussion and it will be taken down.
Where the model looks
The attention/ models also return the last layer's attention, so you can draw maps like these (see Attention maps).
On a film it never saw
None of the models trained on Dune (2021). Across its 65 film-grab stills v5 spreads over all seven sizes.
Labels
The output is 7 logits in this order:
| Index | Label | Short | What fills the frame |
|---|---|---|---|
| 0 | closeUp |
CU | a face |
| 1 | detail |
D | part of a face, an object or text |
| 2 | extremeLongShot |
XLS | a landscape or setting; people tiny or absent |
| 3 | fullShot |
FS | a whole body, head to feet |
| 4 | longShot |
LS | a whole figure, small in its setting |
| 5 | mediumCloseUp |
MCU | head and chest |
| 6 | mediumShot |
MS | from the waist up |
Files
| File | Version | Notes |
|---|---|---|
dinov3_shot_v5_int8.onnx |
v5, int8, 22 MB | recommended: best clean score, used by AiRoll |
mobilenetv3_shot_v6.onnx |
v6, fp32, 16 MB | for video and the browser: MobileNetV3-L distilled from v5, about 50 ms a frame in WebAssembly, 1 point behind v5 |
mobilenetv3_shot_v6.pt |
v6, PyTorch, 17 MB | weights to fine-tune: timm.create_model("mobilenetv3_large_100", num_classes=7), then load_state_dict(torch.load(...)) |
dinov3_shot_v5.onnx |
v5, fp32, 87 MB | |
dinov3_shot_v4_int8.onnx, dinov3_shot_v4.onnx |
v4 | |
dinov3_shot_v3_int8.onnx, dinov3_shot.onnx |
v3 | |
dinov2_shot_v2.onnx |
v2, fp32 | DINOv2 backbone, 14-pixel patches |
attention/*_attn.onnx |
v2 to v5, int8 | same models with a second output: the last attention softmax |
dinov2_shot_tract.onnx |
legacy | 8 classes (incl. ambiguous), made for an early tract WebAssembly demo |
All take pixel_values, float32 [1, 3, 224, 224], and return logits, float32 [1, 7]. Batch size is fixed at 1.
v6 is a convolutional network, so it has no attention variant.
The attention/ files, dinov3_shot_v3_int8.onnx and mobilenetv3_shot_v6.onnx also carry their author, license, label order and preprocessing in
the ONNX metadata (session.get_modelmeta().custom_metadata_map).
Usage
Preprocessing matches training: resize so the short side is 256 px (bicubic), centre-crop 224 × 224, scale to 0–1 and normalise with the ImageNet mean and standard deviation.
import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download
from PIL import Image
LABELS = ["closeUp", "detail", "extremeLongShot", "fullShot", "longShot", "mediumCloseUp", "mediumShot"]
def preprocess(path):
img = Image.open(path).convert("RGB")
s = 256 / min(img.size)
img = img.resize((round(img.width * s), round(img.height * s)), Image.BICUBIC)
left, top = (img.width - 224) // 2, (img.height - 224) // 2
x = np.asarray(img.crop((left, top, left + 224, top + 224)), np.float32) / 255
x = (x - [0.485, 0.456, 0.406]) / [0.229, 0.224, 0.225]
return x.transpose(2, 0, 1)[None].astype(np.float32)
session = ort.InferenceSession(hf_hub_download("szymonrucinski/what-a-shot", "dinov3_shot_v5_int8.onnx"))
logits = session.run(["logits"], {"pixel_values": preprocess("frame.jpg")})[0][0]
probs = np.exp(logits - logits.max())
probs /= probs.sum()
print(LABELS[probs.argmax()], round(float(probs.max()), 3))
In the browser, with onnxruntime-web, the file loads straight from this repo:
const session = await ort.InferenceSession.create(
"https://huggingface.co/szymonrucinski/what-a-shot/resolve/main/dinov3_shot_v5_int8.onnx");
const { logits } = await session.run({ pixel_values: new ort.Tensor("float32", pixels, [1, 3, 224, 224]) });
Attention maps
session = ort.InferenceSession(hf_hub_download("szymonrucinski/what-a-shot", "attention/dinov3_shot_v5_int8_attn.onnx"))
logits, attn = session.run(None, {"pixel_values": preprocess("frame.jpg")})
# attn: [1, 6 heads, 201 tokens, 201 tokens] = CLS, 4 register tokens, 14 x 14 patches
cls_to_patches = attn[0].mean(0)[0, 5:].reshape(14, 14)
# v2 (DINOv2, no registers, 16 x 16 patches): attn[0].mean(0)[0, 1:].reshape(16, 16)
Results
Measured on the human-labelled test frames (859 carry one of the seven sizes). v4 to v6 are scored on the 643 frames none of them trained on:
| Version | Backbone | Trained on | Accuracy | Balanced accuracy |
|---|---|---|---|---|
| v2 | DINOv2-S/14 | 80% of the human frames | 82.3%* | 84.5%* |
| v3 | DINOv3-S/16 | 80% of the human frames + 2,000 auto-labelled | 74.2%* | 77.3%* |
| v4 | DINOv3-S/16 | 53K rows (48K distinct stills) labelled by v3, least confident 6% reviewed | 67.8% | 71.2% |
| v5 | DINOv3-S/16 | the same 53K, least confident 17% reviewed | 69.2% | 72.2% |
| v6 | MobileNetV3-L | distilled from v5 (T = 4, α = 0.8) on the deduplicated stills | 68.3% | 71.9% |
216 of the 859 test frames turned out to be identical to stills in the scraped training rows, so v4 and v5 trained on them; those 216 are left out above. v6 was trained without them (on all 859 it scores 67.6%). v5's epoch was also chosen on the test frames, so its number is slightly optimistic. * Not comparable: v2 and v3 trained on about 80% of the human-labelled frames.
Most of v5's mistakes (81%) are one size off, and medium close-up versus close-up is the hardest boundary. Shot size has soft edges; people disagree there too.
Limitations
- One frame decides: the model sees a single still, centre-cropped to a square, so very wide frames lose their edges.
- Wide shots where people are tiny or absent (landscapes, clouds, machines) often come out as
detail: with no body to measure against, the model falls back on texture. - The training rows contain 5,140 duplicates (the same still filed under two films by a scraping error); labels are unaffected, but those stills count twice.
- Trained on stills from feature films; phone video, animation or sports may behave differently.
- int8 results can differ slightly between CPUs and WebAssembly (logits by about 0.1); the top class rarely changes.
License and credit
These models are released under CC BY-NC 4.0:
- Non-commercial only. Research, teaching and personal projects are fine; commercial use is not allowed.
- Credit is required, including for derivatives. If you use these models or anything fine-tuned, distilled, quantised or otherwise derived from them, name Szymon Ruciński and link huggingface.co/szymonrucinski/what-a-shot in your model card, paper or app, and keep this license for the derived weights.
- v3 to v5 are fine-tuned from Meta's DINOv3 and v2 from DINOv2, so the DINOv3 License and DINOv2's Apache 2.0 terms apply too.
@misc{rucinski2026whatashot,
author = {Ruciński, Szymon},
title = {What a Shot: film shot-size classifiers},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/szymonrucinski/what-a-shot}}
}
Model tree for szymonrucinski/what-a-shot
Base model
facebook/dinov2-small






