Aircraft ID v7

Aircraft ID v7 is a fine-grained image classifier for commercial airliners. Given a single photograph, it predicts the operating airline and the aircraft type as two independent outputs.

Architecture ConvNeXt-Small backbone, two linear classification heads
Parameters 50 M
Input RGB, 768 × 512 (W × H), letterboxed
Airline output 1,026 classes (airline families)
Type output 238 classes (ICAO Doc 8643 type designators)
Formats PyTorch (safetensors), ONNX
Licence CC BY-NC 4.0

Evaluation

Results on a held-out validation set of 20,000 images, single view, no test-time augmentation. The data are split by aircraft registration, so every airframe in the evaluation set is unseen during training.

Task Top-1 Top-5 Mean per-class accuracy Images Classes evaluated
Airline 93.7% 97.2% 81.2% 18,445 1,004
Aircraft type 95.4% 99.1% 89.1% 6,454 233

Top-1 accuracy is weighted towards frequent classes. Mean per-class accuracy weights every class equally and reflects performance on rare airlines and types.

Training data

The dataset contains 1,145,267 photographs of commercial aircraft taken between 2015 and 2026.

Label sets. Each image carries one or both of two labels.

Label set Images Classes Definition
Airline 973,779 1,026 Airline family. Subsidiaries and regional brands are merged into the parent airline.
Aircraft type 352,387 238 ICAO Doc 8643 type designator, verified against an aircraft registration database.

261,415 images carry both labels.

Class balance. Each class is capped at 5,000 images. Classes with fewer than 50 images are excluded.

Split. Images are partitioned by aircraft registration, so all images of a given airframe fall in a single partition.

Partition Images
Train 910,351
Validation 116,145
Test 118,771

Training procedure

  • Initialisation: convnext_small.fb_in22k_ft_in1k (ImageNet-22k pretraining, ImageNet-1k fine-tuning); all layers fine-tuned.
  • Objective: sum of two cross-entropy losses with label smoothing 0.1. An image contributes only to the head for which it has a label.
  • Resolution schedule: progressive, 576 × 384 followed by 768 × 512.
  • Optimisation: AdamW, base learning rate 1e-4 (4× on the classification heads), weight decay 0.05, cosine decay with linear warm-up, mixed precision.
  • Augmentation: aspect-preserving random crop (area scale 0.8–1.0), horizontal flip (p = 0.3), mild colour jitter.
  • Hardware: one NVIDIA RTX 3080 Ti.

Preprocessing

The model expects the same preprocessing used in training:

  1. Apply EXIF orientation and convert to sRGB.
  2. If the longer side exceeds 1600 px, downscale it to 1600 px.
  3. Letterbox to 768 × 512: scale the image to fit while preserving its aspect ratio, centre it, and pad the remainder by replicating the edge pixels.
  4. Normalise with ImageNet mean and standard deviation.

Stretching, centre-cropping or zero-padding changes the input distribution and reduces accuracy.

Usage

pip install torch timm safetensors pillow numpy
python inference.py photo.jpg

inference.py contains the model definition, the preprocessing above and optional horizontal-flip test-time augmentation.

from inference import load, predict

model, airlines, types = load("cuda")
(airline_p, airline_idx), (type_p, type_idx) = predict(model, "photo.jpg", "cuda")
print(airlines[airline_idx[0]], types[type_idx[0]]["code"])

Files

File Description
model.safetensors Model weights, fp32
inference.py Reference implementation of preprocessing and inference
config.json, model_spec.json Input specification, normalisation constants, output dimensions
labels_airline.json Airline class names, index-aligned with the airline output
labels_aircraft_type.json Type designators and names, index-aligned with the type output
onnx/aircraft_v7.onnx ONNX export; normalisation and softmax are included in the graph

Limitations

  • Scope. The model is trained on commercial airliners. General-aviation, business and military aircraft are absent or sparsely represented, and predictions for them are unreliable.
  • Temporal coverage. The training data begin in 2015. Liveries retired before 2015 are not represented.
  • Defunct airlines. Airlines that are no longer operating are not included in the airline label set.
  • Closed label sets. Both outputs are closed-set. An airline or type outside the label sets is assigned to the most similar known class.
  • Closely related variants. Variants within a family (for example A320 and A320neo, or 737-800 and 737 MAX 8) can be confused, usually in close-up or oblique views where the distinguishing features are difficult to resolve.
  • Class imbalance. Accuracy is lower for classes with few training images, as reflected in the mean per-class figures.
  • Image composition. Accuracy degrades when the aircraft occupies a small part of the frame, is heavily occluded, or when several aircraft appear in one image.

Intended use

The model is intended for research and non-commercial use in aircraft recognition. It is not intended for safety-critical or operational decision-making.

Licence

Released under CC BY-NC 4.0. Commercial use is not permitted.

Downloads last month
10
Safetensors
Model size
50.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TerminalAero/aircraft-id-v7

Quantized
(1)
this model