The question is old and simple: how quiet a tone can you still hear? Audiologists answer it with a pure-tone threshold search — play a tone, ask, play it quieter, ask again — across a standard set of frequencies, in a soundproof booth, through calibrated headphones. The phone version keeps the search and loses the booth, and the interesting engineering is entirely in what it does about that loss.

Let's get the framing right first, because the activity itself insists on it: this is a screening estimate through a phone's loudspeaker, not an audiogram. A loudspeaker in a room is not an audiometer in a booth. The result screen says so, and the saved measurement carries its own uncertainty so someone reading it a year later sees the caveat too, not just the curve.

Why it's a protocol, not a capture

Most measurements on the platform are captures: point a sensor at the world, press a button, extract features. A hearing threshold can't be captured, because a single tone tells you nothing. What tells you something is a procedure: an adaptive ladder at each of six audiometric frequencies, walked one band at a time. The tone plays; you press "I heard it" or "nothing" — that's the entire input the test takes; the ladder steps down when you hear and up when you don't, and the level where it oscillates is your threshold for that band.

A procedure has preconditions, and every one of them can fail in a way that must end the run rather than quietly colour the number. That's the design principle, and it shows up three times.

Pin the volume, or the ruler stretches

The first precondition is boring and load-bearing: media volume pinned to maximum for the whole run. Every level the test presents is an attenuation — so many dB below the phone's full output. That's only a ruler if full output means the same thing every time. Pinning the volume is what makes "40 dB down" at trial 60 mean exactly what it meant at trial 1; a volume rocker nudged mid-run would silently re-scale every answer after it.

Measure the room before testing the ears

A booth exists to remove ambient noise. A living room doesn't, so the activity measures the room instead: before any test tone plays, the microphone records the ambient level in each frequency band. The room's noise floor then sets each band's ladder floor — a room that leaves 50 dB of usable range above its noise can measure a threshold perfectly well within that range.

And a band the room is simply too loud for? It's ruled untestable before it's ever presented. Not tested and discarded afterwards — ruled out up front, which is the difference between a four-minute test and a six-minute one that ends by throwing a third of itself away. It also means the screen can tell you which frequencies it's skipping and why, while you can still do something about it: close a window, kill the fridge hum, move rooms, try again.

Refusing beats guessing. The tempting design is to test every band regardless and let noisy conditions widen an invisible error bar. This activity does the opposite: a band that can't be measured properly contributes nothing — not a flagged number, not a shaky estimate, nothing. No band that failed its room check is allowed into the result. A blank cell that tells you why it's blank is worth more than a value you'd have to distrust, because you can act on the blank and you can't un-believe the value.

The loopback: making the speaker confess

The third precondition is the clever one. Phone speakers are wildly non-flat — strong at 2 kHz, weak at 250 Hz, different for every model. So before testing, the phone plays its test bands and listens to itself through its own microphone. This loopback does two jobs: it confirms the speaker actually produced each band that was asked of it (a band the loopback never heard is ruled out, same as a loud room), and it measures the per-device frequency shape — how this handset's output tilts across the bands — anchoring a correction so thresholds at different frequencies are comparable to each other. It's why the test requires a working microphone even though the test itself is pure playback: without the loopback there's no shape correction and no room check, and both are load-bearing for the claim being made.

What the numbers mean — three links, one weak

So how comparable is the result, really? The honest accounting has three links:

That chain makes a run comparable to another person's run on another phone to within a screening instrument's precision — which is the useful claim. But compare against your own earlier runs on the same handset and the weak link disappears: the assumed offset becomes a shared constant that cancels, and the comparison tightens well past ±10 dB. Tracking your own hearing over time on one device is the strongest thing this test can do — the same devices-cancel-in-a-ratio logic that powers the two-phone insulation test, applied across time instead of across a wall.

When a run's anchor can't confirm the reference band, the activity degrades explicitly rather than pretending: it reports plain attenuation instead of estimated hearing level. Less interpretable, still perfectly comparable against another run on the same handset — a smaller true claim instead of a larger false one. And the stored payload keeps the raw attenuation for every band regardless, so if a better anchor arrives later it can re-date old runs instead of only improving new ones.

The small print, kept large

Even a band that ran cleanly may end without a threshold, and each ending gets its own name: still hearing the quietest tone the handset can produce means your threshold is below the floor — you hear better than this phone can measure — while never hearing even the loudest reads as above the ceiling. Neither is folded into a number. Progress, similarly, is counted in bands settled, not trials completed, because an adaptive test doesn't know its trial count in advance — the only honestly monotonic thing to count. And a run abandoned after two taps saves nothing at all; the save button stays disabled until at least one band has genuinely settled.

Use it as what it is: a way to notice that your left ear's high frequencies drifted since spring, or that the new earplugs left a measurable shadow — and a prompt to see a professional when something moves. Its sibling, the vision test, runs the same adaptive-staircase machinery with light instead of sound, on measurably stronger footing. For everything else a phone can measure, start with the tools overview.