Mohua · on-device birdsong recognition for Aotearoa Plate I · Pachycephalidae

Mohua

Mohoua ochrocephala · yellowhead · declining

An iOS 27 app, in development. Every figure on this page is read from the current build.

Woken at daybreak in Tōtaranui by birds he judged “the most melodious wild musick I have ever heard”. Joseph Banks, aboard the Endeavour, 17 January 1770

Two and a half centuries later the chorus is thinner, and this is an instrument for counting what is left of it. A phone listens, and names the birds it is sure enough about - which is fewer than it can hear.

This page is set in the manner of the naturalists' monographs. The plate is genuine: Keulemans, for Buller's A History of the Birds of New Zealand, 1873. Banks himself belongs to an earlier century - he sailed with Cook and died in 1820 - but his ear is the one this instrument is trying to reproduce, so he keeps the epigraph and Buller keeps the printing.

Hand-coloured lithograph of two birds on a branch: a brown and white pōpokotea above, a yellow-headed mohua below.
Pōpokotea & Mohua Mohoua albicilla, Mohoua ochrocephala J. G. Keulemans, 1873
Entry II

The instrument

A general model of the world's birds, cut down to one archipelago

Google DeepMind's Perch 2.0 knows 14,795 classes of sound. New Zealand does not need them. The model was opened, inspected, and modified - four tensors, kept bit-for-bit at the indices that survive - leaving the birds you can actually hear here, and twelve classes of noise so the app can say what is drowning them out.

The slice is not an approximation. At the kept indices the reduced model's logits are identical to the original's; converting to 16-bit for the phone costs an embedding cosine of 0.999928 and no change at all in which bird comes first.

Nothing is filtered. Wind, rain, traffic, aircraft, dogs and voices are detected, not removed - Perch was trained through a plain log-mel frontend on noisy field recordings, so cleaning the audio moves it off-distribution and the artefacts of denoising look like faint birds.

Fig. 1

The slice

215 kept 14,580 discarded

The whole bar is Perch 2.0 - every class it knows. 1.45% of it is New Zealand.

New Zealand birds kept215

From 491 taxa in DOC's threat classification, after merging subspecies Perch cannot separate and dropping the extinct and the vagrant. 52 families.

Conditions kept12

Twelve FSD50K classes collapsing to 9 plain words: wind, rain, traffic, voices, engine, aircraft, thunder, running water, a dog.

Birds with no Perch class at all25

Shipped anyway, marked unhearable. Silence about a bird must mean no class for it, never no bird called.

On disk
27.4megabytes, down from 413. Runs on the phone; nothing is sent anywhere.
Te reo names
161159 of them from the OSNZ Checklist, 2022 - the authority, not a guess.
Entry III

The register of standing

What the Department of Conservation says about the birds this thing can hear

Fig. 2

Threat classification · 190 birds

  • Of concern - declining or worse
  • All other categories
Nationally Critical
13
Nationally Endangered
9
Nationally Vulnerable
21
Declining
16
Nationally Increasing
7
Recovering
5
Relict
14
Naturally Uncommon
10
Data Deficient
1
Migrant
22
Coloniser
10
Not Threatened
27
Introduced and Naturalised
35

59 of 190 are declining or worse. Hearing a kōkako must not look like hearing a blackbird, so the threat status travels with every detection and the interface gives it weight.

The categories are ordinal and the code treats them that way: they can be compared, never averaged. There is deliberately no numeric value on them to tempt an arithmetic mean of a bird's peril.

A further 25 species report no single status. Perch cannot separate the subspecies DOC assessed differently - one class, two verdicts - so the app returns nothing rather than picking the more alarming of the two, and the interface describes the range instead. The first implementation of this got it wrong in an instructive way, and the decision record keeps the wrong version on the page.

Entry IV

The ledger of refusals

The most important page in the book

Perch's own model card says the species logits are uncalibrated and possibly unreliable for rare species, and recommends tuning thresholds on your own data. So 7,489 recordings were scored, and a threshold was fitted per species by cross-validation. 123 species earned one. 90 did not, and the app will not name them.

Fitted
123Median precision 0.80, fitted on a median of 20 clips.
Declined
9060 for want of recordings; 30 because the model is measurably wrong.

A refusal is not a bug being hidden. It is the recorded result of a species whose false positives were counted and found unacceptable - and in most cases the recordings can say precisely which other bird the model is hearing instead.

Fig. 3

Where the recordings came from

iNaturalist
4,722
Xeno-Canto
1,840
NZ Birds Online
867
DOC
60

Mostly Creative Commons; NZ Birds Online's are used for non-commercial research and never redistributed. A bare except: continue once hid 1,814 of these - every AAC file - and still produced a confident-looking table. Decode failures are now counted by reason.

Read the top of that list carefully. Perch has essentially one concept for the brown kiwi: it calls tokoeka “kiwi-nui” at about the same rate it calls kiwi-nui “kiwi-nui”. Kiwi pukupuku, just as congeneric, separates cleanly - so this is a specific failure, not a general one. Tūī and korimako, the two loudest birds in the forest, take each other down together. All 30 are listed in the appendix.

Entry V

The gate that runs all night

A ceiling on everything above - measured at last

Monitor mode runs the model only on the moments an onset gate lets through, because running it continuously would empty the battery by breakfast. That gate is therefore a hard ceiling on what the app can ever name. A call it misses is not a wrong answer - it is no answer, and the interface cannot tell that apart from the bird never having called.

It had never been measured. Nobody has annotated onsets in the harvest, so the model supplies its own ground truth: the five-second window a clip's labelled species scores highest in is, by Perch's account, where the bird is. If the gate does not open there, that bird cannot be named.

Cost is measured on 200 held-out soundscapes. Held out is what makes them usable - measuring on them tuned nothing, and nothing was tuned. The gate's sensitivity was left exactly where it was.

Recall, where selective
94.9%94 of 99 clips on which the gate is demonstrably being selective rather than firing constantly.
Duty cycle
2.9%Of frames, on held-out soundscapes. About 219 inferences a minute.
Species missed entirely
0Of 139. No bird in the corpus was silent to the gate on every one of its clips.
Headline figure
98.7%Discarded. Over all 375 clips the gate also fires on 54% of frames - at that rate it finds any window by coincidence.

That last number is on the page because throwing it away was the finding. A corpus built to answer one question cannot be assumed to answer a different one, and the 54% was first read as a defect before the soundscapes settled it.

Entry VI

No percentages

The one rule the whole thing is built around

Every other bird app renders an uncalibrated logit as “87% confident”. It is not 87% of anything.

So the confidence type has no percentage formatter, and it is impossible to construct a claim without the calibration that backs it. The evidence travels with the assertion, structurally, rather than by anyone remembering to attach it. Non-finite values collapse to zero, never to one: a broken pipeline must not become a confident detection.

Confirmed

Over a threshold fitted for this species on real recordings.

Mohua - threshold 0.0481, fitted on 58 clips, precision 0.78, recall 0.62. The sample size is carried with the claim.

Probable

Over the global default, because this species has no fitted threshold.

The word itself admits the threshold is a default. One of the 90 declined species can never be more than this.

Possible

Heard, below threshold, recorded but not asserted.

Kept because a birder's own ear is better than the model's, and the clip is still there to listen to.

Detections are stored as the raw score, never as the word. When the calibration improves, every record already taken improves with it - rather than a past claim being frozen at a confidence nobody can now justify.