Butterfly and moth identification from wing photos

This repository contains classification heads and a wing cropper for Neotropical butterflies and moths. The frozen image backbone is BioCLIP 2.5-H, downloaded separately. Predictions support identification review and database curation.

The AI Identifier uses the single-photo head through the inference Space. The collection gallery uses a separate attention ensemble for specimens with verified dorsal and ventral photos. Uploaded photos use the corrected v6 crop and feature pipeline. On 27 September 2026 the single-photo head was retrained with 457,091 additional GBIF photographs (research-grade iNaturalist observations and museum specimens, up to 1,000 observations per species), on top of the 558 label corrections from 26 September. A second 27 September release applied 757 vetted taxonomy corrections (Butterflies of America as the authority): duplicate and misspelled species were merged into their valid names, and non-adult and mislabelled training images were removed. The old-to-new name map is taxon_name_map.json. On 29 September 2026 moths were added to the single-photo head: about 8,000 moth species, trained with specimen and light-trap photographs and about 102,000 research-grade iNaturalist moth observations recorded before 12 February 2026.

Files and coverage

File Purpose
head_hier.pt Single-photo cosine head and taxonomy, with 1,024 input dimensions, 15,347 finest-rank outputs, 11,965 species names (about 4,100 butterflies and 8,000 moths), 2,822 genera and 63 families. Finest-rank outputs include named subspecies and species-only labels. Moth species use GBIF backbone names; their tribe and subfamily are placeholders (UNKNOWN_tribe_<genus>, UNKNOWN_subfamily_<genus>). The 27 September butterfly head (7,482 outputs) is at the previous revision of this repository.
taxon_name_map.json Old-to-new species and leaf names for the 27 September taxonomy corrections.
wing_seg_v6.pt Released Wings v6 detector used for the current taxonomy crop.
wing_seg.pt Historical cropper retained for the separate sex-prediction feature path.
region_checklist.json Geographic checklist used by the existing inference workflow.
collection_attention_20260926/ Three attention heads for verified collection photo pairs, standalone model definition, hash manifest and validation report. This does not enable paired uploads in the public Space.

The data cover the major Neotropical butterfly families and about 8,000 moth species, mainly Erebidae, Sphingidae, Geometridae, Saturniidae, Crambidae and Noctuidae. Butterfly sampling is concentrated in Ithomiini. Predictions on sparsely sampled groups need particular care. Higher-rank taxonomy contains incomplete or unresolved entries; output coverage does not establish accuracy for every family or genus.

Single-photo pipeline and field benchmark

For automatic taxonomy prediction, the Space uses the released Wings v6 detector, applies EXIF orientation once, and forms a tight rectangular crop around the resolved wing detections. It keeps the original background inside the rectangle. It places the crop on a grey square with RGB value 128 and resizes to 224 × 224, then extracts and L2-normalizes the frozen BioCLIP image feature. The cosine head predicts all 15,347 leaves, which are summed to species and genus.

The gallery uses the complete leaf distribution. With no location supplied, it applies no geographic prior. A user-selected location still enables the checklist weighting. Manual crop selection remains available. The sex predictor retains its historical image-feature path.

The release pipeline (29 September 2026 head) was evaluated on 4,566 held-out photographs linked to GBIF's iNaturalist Research-grade Observations dataset, covering 629 species in Ecuador, Colombia and Peru. Recorded observation dates and source creation dates follow the BioCLIP 2.5-H February 2026 weight release. Research Grade status, image membership and provider species agreement were checked. This field cohort is separate from the dissected Sanger collection benchmark.

Rank Photographs Top-1 Top-5
Species 4,566 85.74% 96.01%
Genus 4,566 95.16% 98.49%

With the previous classification head and complete probabilities without geography, species Top-1 was 67.46% under the previous preprocessing, 72.67% with corrected v6 preprocessing, and 73.37% after adding a frozen 92-name map. The label-corrected head reached 73.92% without that map. Adding GBIF photographs for every matched species then raised species Top-1 to 85.15% (seed 1701; the three-seed mean on all 4,738 field photographs rose from 72.92% to 85.40%, paired bootstrap 95% interval for the gain from a 100-observation cap to a 1,000-observation cap: 1.67 to 2.94 points). Training photographs were limited to observations recorded before 12 February 2026, so they cannot overlap the test observations; exact image hashes were checked and no training image had a near-identical feature match (cosine 0.995 or above) to a test photograph. After the taxonomy corrections, test labels are scored under the corrected names (the corrections change the recorded names of 586 of the 4,738 field photographs): the corrected head scores 85.74% species and 95.38% genus Top-1 on the GBIF subset, the same accuracy as the pre-correction head scored under the same names (three-seed mean difference on all field photographs -0.14 points, 95% interval -0.45 to 0.20). The gains are 5.21 percentage points from the tested preprocessing package, 0.70 from the name map, 0.55 from the label corrections and 11.23 from the added GBIF photographs. The preprocessing comparison includes detector, crop geometry, padding, normalization and recorded encoder precision; it does not isolate the detector's contribution. The test cohort was examined during several experiments, so these comparisons are exploratory.

Moth evaluation

Moth accuracy was measured on 10,582 research-grade iNaturalist moth photographs from 1,387 observers whose photographs were not used for training (Neotropical records before 12 February 2026). Scores are exploratory: the test photographs come from the same platform as the training photographs, and research-grade identifications are community identifications.

Evaluation Photographs Top-1 Top-5
Species, species in the label space 9,750 85.15% 96.86%
Family, all photographs 10,582 98.38%

Adding moths did not measurably change butterfly accuracy on the 4,566 GBIF field photographs: species Top-1 85.74% (unchanged) and Top-5 96.01% (was 96.23%); three-seed mean difference in species Top-1 -0.04 points, 95% interval -0.36 to 0.26. 0.41% of butterfly field photographs now receive a moth name as the top prediction.

Collection attention model

The model receives separate L2-normalized BioCLIP feature vectors for the dorsal and ventral photos of one specimen. A shared residual projection, view-role embeddings and gated attention combine the available views before the hierarchical cosine head predicts the finest taxon. Training includes missing-view examples. Attention weights describe the contributions of photo features, not pixel-level explanations.

Each of the three checkpoints has its saved calibration temperature. The release averages their complete leaf distributions and applies the previously selected species calibration. It retains each species' conditional distribution over its leaves, so predictions remain consistent across ranks. The gallery then applies its existing Andes/Ecuador geographic prior and sums leaf probabilities to higher ranks. The exact recipe and hashes are in collection_attention_20260926/manifest.json.

The gallery export covers 3,829 verified pairs using the exact cached features bound by training. Twenty historical fallback records and the separate sex predictions are preserved. Gallery predictions include specimens used in training; the held-out results below evaluate separate test specimens.

Paired-photo evaluation

The attention ensemble was evaluated on 314 held-out Sanger pairs. Named-subspecies results include the 296 pairs with a supported recorded named-subspecies label.

Rank Specimens Top-1 Top-5
Named subspecies 296 89.86% 98.65%
Species 314 95.22% 99.04%
Genus 314 98.09% 99.36%

These values are unchanged by the existing geographic prior on this test cohort, although the prior changes probabilities. In the matched experimental comparison, attention improved species Top-1 from 92.04% to 95.22%, a gain of 3.18 percentage points over averaging the single-photo ensemble's predictions across the two views. The paired bootstrap 95% interval was 0.96 to 5.73 points.

A separate comparison with historical five-fold out-of-fold predictions uses only specimens with matching eligible recorded labels. With the geographic prior, species Top-1 increased from 94.67% to 96.60% on 294 pairs, or 1.93 points. Named-subspecies Top-1 increased from 90.05% to 91.01% on 278 pairs, or 0.96 points. Their intervals include zero, and named-subspecies Top-5 decreased from 99.64% to 99.28%. Historical scores average three individually evaluated seeds; the attention score evaluates a three-model ensemble. This comparison is not an isolated architecture ablation. Fold-vocabulary-absent historical truths count as misses.

The test cohort has been inspected during several experiments. These comparisons are exploratory; the intervals condition on the fitted models and do not account for all model-selection uncertainty. They do not measure accuracy on single uploaded photographs.

Previous collection evaluation

The earlier concatenated dorsal/ventral collection method was evaluated across three seeds and five folds, with the side-of-Andes and Ecuador prior:

Rank Specimens Top-1 Top-5
Named subspecies 2,613 87.93% 97.33%
Species 3,355 91.33% 97.91%
Genus 3,806 95.55% 99.26%
Family 3,824 99.32% 99.90%

These denominators differ from the new attention test. Subtracting the two tables does not measure an improvement. Each rank includes specimens with an eligible recorded identification; absent fold-vocabulary labels count as misses.

Single-photo label repair

558 training labels were corrected using captured image-specific source captions, for example Adelpha p phylaca to Adelpha phylaca phylaca. The head was then retrained with the recipe of the previous public head: same 58,165 training images, features, epochs and hyperparameters. Retraining that recipe without the corrections reproduced the previous head (same top prediction on every field photo). The corrections emptied the 92 classes that the frozen name map had redirected, so the map is no longer needed. The released head is seed 1701. Most of the gain is on photos of the corrected taxa; accuracy on other species is unchanged within noise. A retrieval-based classifier and a variant that merged species-only classes into their subspecies were tested and not adopted.

Training data

Taxonomic labels and photographs were compiled from Butterflies of America, Sangay, Noreste, Cotacachi, other Neotropical butterfly databases, and specialist moth resources including the Sphingidae Taxonomic Inventory. The single-photo head also uses 457,091 GBIF photographs (research-grade iNaturalist observations and museum specimen images recorded before 12 February 2026, species-level labels, up to 1,000 observations per species); they are used only for training and are not redistributed. The augmented attention experiment used a smaller verified GBIF set. Dataset membership, image provenance and split identity were frozen for the reported experiments.

License

CC-BY-NC-4.0, attribution and non-commercial use. Please credit this work and respect the terms of the underlying image and taxonomic sources. The BioCLIP backbone has its own upstream license and is not redistributed here.

This is an AI suggestion, not a definitive identification.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using fr4nzzch/butterfly-id-classifier 1