Hearing the room
Eight situations where the implant delivers sound but not meaning. All of them are the same
underlying problem, too little spectral resolution to separate one voice from the rest, and
all of them are attacked the same way: improve the ratio before the signal reaches the implant,
and put what is left in the visual field.
01
The restaurant
Multi-talker babble, hard surfaces, no visual line to most speakers
≈ −5 dB SNR
Off-frame prototype
What breaks
This is the hardest ordinary environment there is, and the one implant recipients name
first when asked what they gave up. Babble is the worst masker because it is speech,
it occupies the same spectrum and the same modulation rates as the target, so no amount
of amplification improves the ratio.
A recipient needs roughly 10–15 dB more favourable SNR than a normal-hearing listener to
reach comparable sentence recognition. A restaurant sits about 15 dB below that line.
Example
Four people at a table. Two side conversations. Someone at the next table laughs.
Your dinner companion asks a question and you get the prosody, the timing and the fact
that it was a question, but not the words. You say "sorry?" for the third time and
then stop asking, and spend the rest of the meal nodding at the right intervals.
Leo: the array steers toward whoever you are facing, suppresses the flanking
tables, and the caption line carries the words the beamformer could not fully clean up.
What it will not fix: a beamformer improves the ratio in the direction you
point it. Turn your head to the person on your left and the person on your right gets harder,
not easier. Group conversation across a round table remains genuinely difficult, and we do not
have a good answer for it yet.
02
The classroom
One distant talker, long reverberation, unpredictable turn-taking
≈ 0 dB SNR
Off-frame prototype
What breaks
Distance and reverberation together. A teacher eight metres away arrives with a poor
direct-to-reverberant ratio, and reverberation smears exactly the temporal envelope cues
that implant users depend on most, the cues Shannon's work showed carry most of the
intelligibility when spectral detail is scarce.
Group work is worse than lecture: the talker changes every few seconds and there is no
podium to point at.
Example
A student follows the lecture adequately. Then the class splits into groups of five,
thirty other students start talking, and the room becomes unusable. They spend the
session copying a neighbour's notes and reconstructing the discussion afterwards.
Leo: in lecture mode the array locks the podium and holds it through head
movement. In group work, captions with speaker attribution mean a glance identifies who
is talking without having to find them first.
What it will not fix: a room-mounted or lapel microphone on the teacher,
streamed directly, still beats anything a head-worn array can do at eight metres. Where such a
system exists, Leo should be complementary to it, not a replacement for it, and we would say
so to a school considering the trade.
03
Meetings and open-plan offices
Rapid turn-taking, overlapping speech, high stakes for missing a name
≈ +5 dB SNR
Off-frame prototype
What breaks
The SNR is survivable; the structure is not. Meetings turn over quickly, people
interrupt, and the cost of missing four words is not embarrassment but a wrong commitment.
Attribution matters as much as content: knowing a sentence was said is useless if you do
not know who said it.
Example
Six people, one speakerphone, two of them talking over each other. You catch "…can you
own that by Thursday?" and have no idea whether it was aimed at you.
Leo: this is the case the spatial map was built for. Each caption line is
tagged to the direction it arrived from, so the display can say who, not just
what, the single highest-value output the visual layer produces.
What it will not fix: remote participants on a speakerphone are a single
mono source from one point in the room. Spatial attribution cannot separate three people
dialled in through the same speaker, and no microphone array ever will.
04
The family dinner table
Several people you love, all talking at once, every week
≈ 0 dB SNR
Off-frame prototype
What breaks
Acoustically this is the restaurant. Socially it is not. Withdrawal at a family table is
cumulative and it is the mechanism by which hearing loss turns into isolation, the
outcome the WHO burden figures are ultimately counting.
Cross-talk is constant, nobody waits their turn, and the people involved are the least
likely to be asked to repeat themselves for the tenth time.
Example
Three conversations at one table. A grandchild says something quietly, everyone laughs,
and you laugh a beat late because you are reading the room rather than hearing it.
The joke is gone before anyone could repeat it.
Leo: captions hold the last few lines, so a beat-late glance recovers what was
said instead of losing it. Attribution means you know which grandchild.
What it will not fix: reading a joke is not the same as hearing it, and a
caption arriving 120 ms late is still late. We think a recoverable conversation beats a lost
one. We do not think it is the same thing.
05
Counters, pharmacies and clinics
Glass barriers, masks, background machinery, safety-critical content
≈ +5 dB SNR
Off-frame prototype
What breaks
Two compounding problems. Barriers and masks attenuate high frequencies, where much of
the consonant information lives. And a mask removes the mouth, which for an implant user
is not a minor loss. Rouger and colleagues showed recipients are unusually strong
lipreaders and multisensory integrators; covering the face takes away a channel they
rely on more than a hearing person would.
Example
A pharmacist behind glass, masked, explaining a dosing change. You hear a muffled
sentence containing a number. You are not certain whether it was "one to two" or
"one or two", and the difference matters.
Leo: the array steers hard at the counter and suppresses the room. The caption
carries the number, which is the part you cannot afford to guess.
What it will not fix: transcription is not perfect and never will be. For
anything safety-critical, a dose, a dosage interval, a diagnosis, the right behaviour is to
ask for it in writing, and we will not build a product that discourages that.
06
Transit and public address
Vast reverberant spaces, compressed PA audio, information you cannot miss
Reverberant
Planned
What breaks
Station and airport announcements are heavily compressed, played through distributed
speakers in enormous reverberant volumes, and arrive from every direction at once. There
is no single source to steer at, the room is the source. Directional gain, the
main tool for every other case on this page, does almost nothing here.
Example
A platform change is announced. Everyone around you picks up their bags and moves. You
watched them react before you understood the announcement, and now you are following a
crowd without knowing where it is going.
Leo: the caption path matters far more than the acoustic path here. Recognition
is run against the diffuse field and rendered as text, because text survives reverberation
in a way that audio does not.
What it will not fix: honestly, this is our weakest case. Heavy compression
plus multi-second reverberation degrades recognition badly, and we would rather say that here
than have someone buy the device for this and be let down.
07
Television and shared media
Compressed dialogue under a music bed, competing with room noise
≈ +10 dB SNR
Off-frame prototype
What breaks
Modern mixes bury dialogue under score and effects, which is difficult for hearing
listeners and disqualifying for implant users. Broadcast captions exist and are good, but
they are on the screen, so you cannot look away, and you cannot use them for the person
on the sofa talking to you about what just happened.
Example
Subtitles are on. Someone asks you a question during a scene. You now have to choose
between the subtitle and the question, and you lose whichever you did not pick.
Leo: the room and the screen are different sources in the spatial map, so the
person beside you can be attributed and captioned separately from the television.
What it will not fix: for the programme itself, broadcast subtitles are
authored, corrected and better than live recognition. Use them. Leo's value here is the sofa,
not the screen.
08
Driving and passengers
Broadband road noise, a talker you must not look at
≈ +5 dB SNR
Planned
What breaks
A passenger sits beside and slightly behind you, and every visual cue an implant user
leans on is unavailable because your eyes belong to the road. This is the one common
situation where the visual channel, the whole basis of the product, is off the table
by law and by common sense.
Example
Your passenger gives directions at 70 km/h with the windows cracked. You either take
your eyes off the road to lipread, or you guess at the turn.
Leo: the array steers to the passenger seat and suppresses road noise, which
is broadband and diffuse and therefore exactly what a beamformer is good at. Captions are
suppressed while driving.
What it will not fix: we will not put a caption in a driver's field of view.
The acoustic path has to carry this case alone, which makes it the one use case where Leo is
strictly a microphone and not a display.