Nocturne v1.2d (Teacher) β€” the first Nocturne model to beat v1 on the frozen test

Audio Spectrogram Transformer, ~86M parameters, 2,196 output columns. This is the first candidate in the Nocturne line to clear our pre-committed release gate against the v1 teacher. It got there by fixing two supervision bugs, not by adding data β€” which is the most useful thing in this card.

Read the limitations section before deploying it. On an unseen recording site this model is substantially worse than v1, and our much smaller distilled student beats it outright on the very test it was released for. Both are stated with numbers below.

Results on the frozen harness

Identical 13,710-clip held-out test set, per-class thresholds calibrated on validation, scored over the 2,182 name-aligned core classes. Rows are identical across all models by construction, so these columns are directly comparable.

model params macro-F1 (calibrated) mAP
v1 teacher (baseline) 86M 0.148766 0.149849
v1.2d (this model) 86M 0.151051 0.154110
v1.2 (weak labels) 86M 0.126 0.126
nocturne-v1-mini (student) 9.4M 0.154198 0.160848

Deltas against v1: +0.002286 macro-F1, +0.004262 mAP. Both exceed the noninferiority margins of 0.00169 and 0.00152, which were derived from a paired bootstrap over the test rows and written into the run's ship condition before this model was scored. That ordering is the point: the margin was not chosen to let the candidate through.

What actually produced the gain

Three earlier runs isolate it, and the answer is not more data.

run change calibrated macro-F1
v1.2 added AnuraSet with weak recording-level labels 0.126 β€” regression
v1.2b strong labels, medium-quality subset 0.1476 β€” parity at best
v1.2c strong labels + fixed annotation parser 0.1474 β€” no gain from the parser alone
v1.2d parser fix + multi-hot targets + mAP checkpoint selection 0.1511

Two bugs, both in supervision rather than architecture:

  1. The annotation parser dropped 56% of AnuraSet annotations. The suffix on a strong-label column is call quality (_L/_M/_H), not sex. Our parser assumed sex and silently discarded everything it did not recognise.
  2. Targets were one-hot when the task is multi-label. About 37% of strong-label rows were the same clip under a different species. One-hot training presented those clips repeatedly with contradictory negatives. Collapsing to multi-hot took training rows from 137,068 to 86,703 β€” fewer rows, better model.

Checkpoint selection mattered as much. Best epoch was 18 of 40, chosen by validation mAP. Under the previous F1@0.5 rule this run would have shipped a much later and worse checkpoint.

Limitations β€” please read these

It is worse than v1 at a recording site it has never heard. On a held-out AnuraSet site, top-1 accuracy is 0.138 against v1's 0.544. Adding an anuran corpus did not buy site generalisation; the v1.2 family learned site signatures. If you are deploying to a new field location with no local validation data, use nocturne-v1-teacher instead. This diagnostic covers 2,209 rows and only 5 scorable classes, so treat it as a warning signal rather than a precise measurement β€” but the direction has been consistent across every model in this family.

Our 9.4M student beats it. nocturne-v1-mini scores higher on the same frozen test at roughly a ninth of the parameters and about ten times the speed. If you want the best numbers on this benchmark, or anything running at the edge, take the mini. This model is published because it is the first teacher to clear the gate and because the result behind it is worth having on the record, not because it is the best model we have.

The vocabulary contains near-duplicate entries. 2,196 columns cover about 1,896 unique species; some appear both as Genus species and Genus_species from differing source conventions. Scores can split across the pair. The 2,182-class core set used for evaluation is name-aligned to handle this.

Not for bird identification. Use BirdNET or Perch. Not for legal or conservation decisions without field verification. Coverage is biased toward temperate zones and well-recorded taxa.

Usage

from model import load_nocturne, predict_file
model, vocab, thresholds = load_nocturne(".")     # strict load, transformers-version aware
print(predict_file(model, "clip.wav", vocab, thresholds, top_k=5))

model.py remaps parameter names across transformers versions and then loads strictly. These weights were saved under transformers β‰₯5.16, which renamed every AST attention parameter. On an older build the names will not match, and loading with strict=False appears to succeed while leaving the entire backbone at its AudioSet initialisation β€” the model then returns confident nonsense. Verified loading cleanly on transformers 4.57.6 and 5.16.1. If it raises, install a matching transformers rather than relaxing the check.

Files

model.safetensors (verified bit-identical to the training checkpoint before upload) Β· vocab.json Β· thresholds.json (per-class, calibrated on validation) Β· eval_report.json (the full frozen-harness report behind the numbers above) Β· training_config.yaml (including the ship condition) Β· config.json

Hosted demo

The hosted demo and API were retired on 2026-09-27; run the checkpoint locally as shown under Usage.

Licence: weights CC-BY-4.0, code Apache-2.0.

vs BirdNET on non-bird taxa (release headline)

Same protocol as v1's card: 1,000 randomly sampled non-bird clips from the iNat Sounds 2024 test split (never trained on), BirdNET v2.4 via birdnetlib (min_conf floor as its author recommends) and this model given identical inputs; metric is top-1 species accuracy. Run 2026-09-23 on a DGX Spark (CPU inference), harness soundscape/birdnet_benchmark.py built with this checkpoint's 2,196-class vocab.

Model Top-1 on non-bird clips (n=1,000)
BirdNET (v2.4, bird-focused) 6.7%
Nocturne v1.2d teacher 73.3%

~11Γ— lift. v1.1 measured 76.7% vs 6.7% on a 300-clip sample; the two samples differ in size and draw, so treat 73.3 and 76.7 as the same finding, not a regression. BirdNET's vocabulary is bird-only, so its non-bird accuracy is expected to be near zero β€” the point of this table is that Nocturne fills that gap, not that BirdNET is bad at birds.

Research page

The write-up and every Nocturne release, in one place: https://runstratus.com/research/nocturne-v1-non-bird-bioacoustics

Downloads last month
63
Safetensors
Model size
87.9M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support