What a Shot: film shot-size classifiers

Small vision transformers that read a film frame's shot size, from extreme long shot to detail, exported to ONNX. They run on CPU or in the browser (ONNX Runtime Web); no PyTorch needed.

  • Try them: What a Shot tags movie trailers live in the browser with v6, and shows every version's prediction and attention on human-labelled test frames.
  • See them used: AiRoll cuts clips to the beat of a song and labels each shot with dinov3_shot_v5_int8.onnx.
  • Data: types-of-film-shots: 53,449 labelled training rows (48,309 distinct stills from about 1,000 films) and 863 human-labelled test frames.
  • License: CC BY-NC 4.0: non-commercial use only, and credit is required, also for fine-tuned models.

Seven shot sizes, from extreme long shot to detail

Live on movie trailers

v6 tagging The Dark Knight trailer live in the browser, with the shot-size timeline

In the What a Shot Space v6 runs on every frame it can while a trailer plays and tags the shot size in the corner; the timeline below draws its answers over v5's. v6 is, a MobileNetV3-L distilled from v5, running in real time in the browser: about 50 ms a frame (about 20 frames a second) on a desktop CPU, in WebAssembly on one thread.

Disclaimer: the trailers are shown for educational and research purposes only, to demonstrate the classifier. I do not own them: all rights belong to their studios and publishers. If you hold the rights to one and want it removed, open a discussion and it will be taken down.

Where the model looks

v5 on one test frame per shot size, with its last-layer CLS attention

The attention/ models also return the last layer's attention, so you can draw maps like these (see Attention maps).

On a film it never saw

Dune (2021): one frame v5 labelled as each shot size, with its attention

None of the models trained on Dune (2021). Across its 65 film-grab stills v5 spreads over all seven sizes.

Where v5 goes wrong: wide shots with tiny or no people labelled as detail

Labels

The output is 7 logits in this order:

Index Label Short What fills the frame
0 closeUp CU a face
1 detail D part of a face, an object or text
2 extremeLongShot XLS a landscape or setting; people tiny or absent
3 fullShot FS a whole body, head to feet
4 longShot LS a whole figure, small in its setting
5 mediumCloseUp MCU head and chest
6 mediumShot MS from the waist up

Files

File Version Notes
dinov3_shot_v5_int8.onnx v5, int8, 22 MB recommended: best clean score, used by AiRoll
mobilenetv3_shot_v6.onnx v6, fp32, 16 MB for video and the browser: MobileNetV3-L distilled from v5, about 50 ms a frame in WebAssembly, 1 point behind v5
mobilenetv3_shot_v6.pt v6, PyTorch, 17 MB weights to fine-tune: timm.create_model("mobilenetv3_large_100", num_classes=7), then load_state_dict(torch.load(...))
dinov3_shot_v5.onnx v5, fp32, 87 MB
dinov3_shot_v4_int8.onnx, dinov3_shot_v4.onnx v4
dinov3_shot_v3_int8.onnx, dinov3_shot.onnx v3
dinov2_shot_v2.onnx v2, fp32 DINOv2 backbone, 14-pixel patches
attention/*_attn.onnx v2 to v5, int8 same models with a second output: the last attention softmax
dinov2_shot_tract.onnx legacy 8 classes (incl. ambiguous), made for an early tract WebAssembly demo

All take pixel_values, float32 [1, 3, 224, 224], and return logits, float32 [1, 7]. Batch size is fixed at 1. v6 is a convolutional network, so it has no attention variant. The attention/ files, dinov3_shot_v3_int8.onnx and mobilenetv3_shot_v6.onnx also carry their author, license, label order and preprocessing in the ONNX metadata (session.get_modelmeta().custom_metadata_map).

Usage

Preprocessing matches training: resize so the short side is 256 px (bicubic), centre-crop 224 × 224, scale to 0–1 and normalise with the ImageNet mean and standard deviation.

import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download
from PIL import Image

LABELS = ["closeUp", "detail", "extremeLongShot", "fullShot", "longShot", "mediumCloseUp", "mediumShot"]

def preprocess(path):
    img = Image.open(path).convert("RGB")
    s = 256 / min(img.size)
    img = img.resize((round(img.width * s), round(img.height * s)), Image.BICUBIC)
    left, top = (img.width - 224) // 2, (img.height - 224) // 2
    x = np.asarray(img.crop((left, top, left + 224, top + 224)), np.float32) / 255
    x = (x - [0.485, 0.456, 0.406]) / [0.229, 0.224, 0.225]
    return x.transpose(2, 0, 1)[None].astype(np.float32)

session = ort.InferenceSession(hf_hub_download("szymonrucinski/what-a-shot", "dinov3_shot_v5_int8.onnx"))
logits = session.run(["logits"], {"pixel_values": preprocess("frame.jpg")})[0][0]
probs = np.exp(logits - logits.max())
probs /= probs.sum()
print(LABELS[probs.argmax()], round(float(probs.max()), 3))

In the browser, with onnxruntime-web, the file loads straight from this repo:

const session = await ort.InferenceSession.create(
  "https://huggingface.co/szymonrucinski/what-a-shot/resolve/main/dinov3_shot_v5_int8.onnx");
const { logits } = await session.run({ pixel_values: new ort.Tensor("float32", pixels, [1, 3, 224, 224]) });

Attention maps

session = ort.InferenceSession(hf_hub_download("szymonrucinski/what-a-shot", "attention/dinov3_shot_v5_int8_attn.onnx"))
logits, attn = session.run(None, {"pixel_values": preprocess("frame.jpg")})
# attn: [1, 6 heads, 201 tokens, 201 tokens] = CLS, 4 register tokens, 14 x 14 patches
cls_to_patches = attn[0].mean(0)[0, 5:].reshape(14, 14)
# v2 (DINOv2, no registers, 16 x 16 patches): attn[0].mean(0)[0, 1:].reshape(16, 16)

Results

Accuracy of each version on 859 human-labelled frames

Measured on the human-labelled test frames (859 carry one of the seven sizes). v4 to v6 are scored on the 643 frames none of them trained on:

Version Backbone Trained on Accuracy Balanced accuracy
v2 DINOv2-S/14 80% of the human frames 82.3%* 84.5%*
v3 DINOv3-S/16 80% of the human frames + 2,000 auto-labelled 74.2%* 77.3%*
v4 DINOv3-S/16 53K rows (48K distinct stills) labelled by v3, least confident 6% reviewed 67.8% 71.2%
v5 DINOv3-S/16 the same 53K, least confident 17% reviewed 69.2% 72.2%
v6 MobileNetV3-L distilled from v5 (T = 4, α = 0.8) on the deduplicated stills 68.3% 71.9%

216 of the 859 test frames turned out to be identical to stills in the scraped training rows, so v4 and v5 trained on them; those 216 are left out above. v6 was trained without them (on all 859 it scores 67.6%). v5's epoch was also chosen on the test frames, so its number is slightly optimistic. * Not comparable: v2 and v3 trained on about 80% of the human-labelled frames.

v5 confusion matrix on the 643 clean test frames

Most of v5's mistakes (81%) are one size off, and medium close-up versus close-up is the hardest boundary. Shot size has soft edges; people disagree there too.

Limitations

  • One frame decides: the model sees a single still, centre-cropped to a square, so very wide frames lose their edges.
  • Wide shots where people are tiny or absent (landscapes, clouds, machines) often come out as detail: with no body to measure against, the model falls back on texture.
  • The training rows contain 5,140 duplicates (the same still filed under two films by a scraping error); labels are unaffected, but those stills count twice.
  • Trained on stills from feature films; phone video, animation or sports may behave differently.
  • int8 results can differ slightly between CPUs and WebAssembly (logits by about 0.1); the top class rarely changes.

The What a Shot Space

License and credit

These models are released under CC BY-NC 4.0:

  • Non-commercial only. Research, teaching and personal projects are fine; commercial use is not allowed.
  • Credit is required, including for derivatives. If you use these models or anything fine-tuned, distilled, quantised or otherwise derived from them, name Szymon Ruciński and link huggingface.co/szymonrucinski/what-a-shot in your model card, paper or app, and keep this license for the derived weights.
  • v3 to v5 are fine-tuned from Meta's DINOv3 and v2 from DINOv2, so the DINOv3 License and DINOv2's Apache 2.0 terms apply too.
@misc{rucinski2026whatashot,
  author       = {Ruciński, Szymon},
  title        = {What a Shot: film shot-size classifiers},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/szymonrucinski/what-a-shot}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for szymonrucinski/what-a-shot

Quantized
(8)
this model

Dataset used to train szymonrucinski/what-a-shot

Space using szymonrucinski/what-a-shot 1