How we measure, so you can check us
Every number the app shows was measured on recordings that people, not this app, labelled. This page is the audit trail: what was measured, against what, and what the numbers do and do not mean. Back to the app.
The rules the numbers live under
Refusing beats guessing. When the app is not confident, it says so and the clip costs you nothing. A call earns its phrase only when, in testing, it is right about 6 times in 7 when it speaks. That bar does not bend for a launch, a feature, or a good week.
Accuracy is never for sale. Free and paid analyse identically. Paying buys time and convenience, never a better ear.
Numbers only move when re-measured. Nothing on this page updates because marketing wanted it to.
What the app claims today, and the test behind each claim
| claim | measured | against |
|---|---|---|
| "That was a horse" | 94% right when it speaks (95% CI 85–98%) |
148 randomly drawn recordings labelled by ear, drawn blind and held out of every tuning decision. Earlier this read 96%/55% from a 44-clip curated research set; the numbers now come from the kind of clips the app actually meets, and the random draw is the only frame this page will quote from now on. |
| Hears a horse sound at all | 71% of real sounds (CI 61–79%) |
same 148 recordings |
| "That was a whinny" | 65% right when it speaks (CI 49–78%), catches 92% (CI 76–98%) |
same 148 recordings. The earlier 89%/73% pair was measured on the curated research set and is superseded by the random draw. |
| Nicker recognizer, in training, not yet speaking | no publishable score yet | The figures previously shown here (82% right, catches 76%) were measured against labels taken from the uploaders' own video titles. When people listened to a random sample of those recordings by ear, most of the clips titled "nicker" were not nickers, and several were human voices, so the labels were not a ruler and the numbers they produced are withdrawn. Rebuilding on recordings people have confirmed by ear; nothing goes back in this row until enough of those exist. |
The ranges are wide because the samples are small, and they are shown because a number without its range is a guess in a suit. Missing sounds is normal: leave it listening and repeated calls make the odds add up.
What the numbers mean
Precision means: when the app speaks, how often is it right. Recall means: of the real sounds, how many did it catch. Every number above is measured against recordings that people, not this app, labelled, and the app is never scored on a recording it learned from. Both numbers have to clear the bar before a phrase ships, because a wrong word in your horse's mouth is worse than none.
It gets better as the community shares sounds. Every shared clip is reviewed by a person before it teaches anything, and the reviewer's verdict, not the model's opinion, becomes the label.
What we will not claim
The app never says what a horse feels or wants; it names calls and offers each call's documented function as a phrase, the way a phrasebook renders one language in another. Trends are counts and arithmetic with the observation time attached, never a diagnosis. And nothing here is a substitute for a veterinarian.