Leo is a pair of glasses for people with cochlear implants. In a loud room, it
raises the one voice you are trying to follow above the noise around it. For Deaf signers it works the
other way, and speaks their signing aloud to people who don't sign.
The hardware
An engineering model you can turn in your hands.
To hear which direction a voice comes from, you need microphones spread across the head.
Glasses already sit there. Here is where each part goes on the DK-1 frame.
Leo DK-1 · engineering model, rendered live
Drag to rotate
Drag to rotate. The frame is generated from the DK-1 design targets and rendered in your browser:
titanium rims and temples, prescription-compatible lenses, six microphone ports across the brow bar
and temples for aperture, two sign cameras under the bridge for the outward channel, and the
waveguide on the right lens. Compute and radio are split across the temples for thermal and mass
balance. To be explicit: this frame has not been built. It is the layout our software is being
written against, not a photograph of something on a desk.
What Leo is
AI glasses for cochlear implant recipients.
Six microphones pull one voice out of a crowded room
Streams only that voice into the implant you already own
Rides the accessory pathway your manufacturer already ships
A second channel reads sign and speaks it aloud
All of it on the frame, so no conversation leaves the device
Where we actually arePre-hardware
The software runs. The frame does not exist yet.
Separation and sign models run on a laptop, on recorded audio
The DK-1 frame has not been built
No study has been run and no clinical claim is made
Every Leo figure below is a design target, labelled as one
What we are raising forThe ask
A whole Leo, on a face, working.
Six microphones on a built frame
Separation and sign both running on-frame
Inside the latency budget, on its own battery
Coupled into a real implant
Worn by recipients, who tell us whether it worked
Why we publish it this way. Assistive hearing technology has a long history of
launching on bench numbers that did not survive a real room. We put the literature we have to beat
on our own site before publishing anything of our own, and we would rather lose a meeting for being
early than win one for being vague. Everything below is either a citation or a labelled target.
The idea
A group conversation leaves people behind in two directions.
You get left out when you can't pull one voice out of the noise, and when the room can't
understand the voice you have. For decades the first half was a hardware problem with no good
fix. What changed is recent: AI can now pull overlapping voices apart in real time and label
who said what, on a chip small enough to wear
[15][16][17].
Leo puts that on your face, and pairs it with a channel that runs the other way.
The room → youInward
Six microphones and an on-frame model split the room into separate voices in real time,
lock onto the person you're attending to, and caption each line tagged to who said
it, so a table of people talking over each other becomes one thread you can actually follow.
For a cochlear-implant wearer, whose hearing plateaus at a handful of usable channels, that
ratio is the only lever left. How the acoustic stack works →
You → the roomOutward
Two cameras under the bridge watch the wearer's signing space; an on-device model reads the
hands and the frame speaks the words aloud. A Deaf signer stops needing a phone, a notepad or
an interpreter just to be understood by someone who doesn't sign.
Everything is read on-frame and the video is discarded; the wearer picks the voice and
confirms before anything is spoken. How sign → speech works →
The whole product is one sentence: everyone in the conversation gets to stay in
it. The rest of this page is the evidence, the engineering, and the honest limits
behind that claim.
Why glasses
The ear is already spent. The eyes are usually next.
Every other company building hearing AI is competing for space inside the ear. For the people we
build for, that space is gone, and it was gone before anyone shipped an earbud. What is left is
the face, and it turns out to be the surface our users were most likely to be wearing something
on anyway.
The ear is occupied
A cochlear implant recipient wears a sound processor on the ear and against the side of the
head, every waking hour. That is not a preference and it is not coming off. Any product shaped
like an earbud is structurally locked out of this user, not because the software is weak but
because the geometry is taken.
Glasses are the only unoccupied surface on the head, and they are the only one with the
sightlines this problem needs. An array on the temples has aperture the ear cannot offer, the
bridge looks where the wearer looks, and the same mount that hears the talker can see the
talker's mouth and the signing space.
Deafness and vision loss travel together
Vision problems affect somewhere between 44 and 65 percent of deaf and hard-of-hearing people,
against 17 to 30 percent of normal-hearing people [26]. Roughly double,
and it is not a coincidence.
The retina and the cochlea differentiate from the same layer in weeks six and seven of
gestation [26]. What damages one frequently damages the other, whether
genetic, hypoxic or viral. Usher syndrome, the leading inherited cause of deafblindness at 4 to
17 per 100,000, is the clearest case, and those patients are routinely implanted
[27].
Vision problems are about twice as common in deaf and hard-of-hearing people
Reported prevalence ranges across the literature. Each bar is the range as stated, not an
interpolated estimate.
Prevalence ranges as reported in a study of refractive error, amblyopia, strabismus and low
vision among deaf and hearing-impaired students, which also states the shared embryonic origin
of retina and cochlea [26]. The underlying
studies weight toward children and adolescents; we have not found an equivalent figure measured
specifically in adult implant recipients, and we are not claiming one.
And it compounds with age
Combined hearing and vision loss climbs steeply in later life: about 1.5% of people aged 65 to
74, 2.6% at 75 to 84, and 10.8% past 85 [28]. The same age band where
hearing loss concentrates is the band where the second sense starts going too.
A device that sits on the face can carry optical correction and audio on one mount. We are not
building display features for their own sake, but the person we are designing for is more likely
than not to need a lens in front of their eye regardless of what we do.
Why this is the adoption argumentThe bet
Only 16% of US adults aged 20 to 69 who could benefit from a hearing aid have ever used one
[19]. That number, not the engineering, is what has killed assistive
hearing hardware for forty years. People do not decline these devices because they do not work.
They decline them because of what wearing one means.
So we are not asking anyone to put a new device on their face. Something close to half of them
already have one there. We are asking them to change frames, and that is the lowest-friction
adoption story available in this category.
Capabilities
What the glasses do, and how far along each part is.
Everything below runs on the glasses themselves. Nothing you say or hear is sent to a server.
Six separate models share one chip on the frame, and the whole path from sound in your ear to
text in your eye is built to stay under 120 ms. The badge on each one says how finished it is,
not what it does in a clinic.
6on-frame neural models, one shared 4.2 TOPS NPU
<120mswavefront to stimulation, inside the binding window [7]
10–15dBthe noise deficit implant users carry, and the number Leo exists to move [3]
0bytes of audio, transcript or video that leave the device
Off-frame prototype
Real-time voice separation
A causal, streaming separation network pulls the talker you are attending to out of overlapping
speech and diffuse babble, on-frame, with lookahead capped at 32 ms. It descends from the line of
work that took source separation from an offline trick to a real-time one [15][16].
Planned
Speaker memory
Enroll the voices that matter, a partner, a child, a manager, in a few seconds each. Leo learns
their embeddings and prioritizes them across rooms and days, and tags every caption with who is
speaking so a crowded table reads as one attributable thread.
Planned
Gaze-steered attention
Gaze direction, head orientation and streaming diarization fuse into a single target estimate,
so Leo locks onto the person you are looking at without you touching a thing. When confidence
drops it widens rather than guesses, because a wrong lock is far more disorienting than a broad one.
Research
Live translation
A speech-translation model renders a talker in another language into the wearer's language,
captioned in the display and optionally spoken. On-device and scoped to conversational domains,
we are not claiming open-domain, unbounded translation, and we would distrust anyone who did.
Planned
Spatial captions
Transcribed speech lands in the lower periphery of a monocular waveguide, each line tagged to
the direction it arrived from and never placed over a talker's mouth. Text and sound are time
aligned to fall inside the same audiovisual binding window, by design, not by luck.
Research
Sign → speech
Two cameras under the bridge read the wearer's signing space, an on-device pose model reads the
hands, and the frame speaks. It gives a Deaf signer a voice with people who do not sign. It is the
most valuable thing we could build and the least finished. Read the honest limits →
These stack on two independent levers, both established in the literature: directional
gain from a head-wide microphone array, and the 10–15 dB effective boost of seeing the talker
[4]. Leo is an attempt to put both on one face at
once, which is exactly the deficit the implant leaves behind [3].
The problem
The implant solved audibility. It did not solve the cocktail party.
A modern cochlear implant restores access to sound with remarkable reliability. What it does not
restore is the ability to pull one voice out of a room full of them. That failure is not a defect
in the device, it is a consequence of what an electrode array can physically do inside a cochlea,
and it has been stable in the literature for three decades.
1.5B
people live with some degree of hearing loss; 430 million require rehabilitation today.
The population is growing faster than the response
People living with hearing loss worldwide, today versus the WHO's 2050 projection.
World Health Organization, Deafness and hearing loss fact sheet, drawing on the
World Report on Hearing (2021). "Rehabilitation" counts moderate or higher loss in
the better ear. [1]
Where the burden actually sits
Share of people with disabling hearing loss living in low- and middle-income countries.
What the hardware provides, against what listeners can actually use. Each bar is the
range reported in the cited work, not an interpolated curve.
Asymptote of 8–10 channels: Berg et al., JASA 147(5), 2020, in 18 adult recipients on
CNC words and AzBio sentences [11]. Continued
gains to 20 channels for normal-hearing listeners on vocoded speech: Friesen et al., 2001
[3]. Electrode counts are the commercial range
across current manufacturers.
Why channels collapse
An implant array sits in a conductive salt bath. Current injected at one electrode spreads to
neighbouring neural populations, so physically adjacent contacts excite overlapping regions of
the spiral ganglion. Adding electrodes does not add independent channels; it adds correlated ones.
Friesen and colleagues showed the consequence directly: normal-hearing listeners given
noise-vocoded speech keep improving as channel count rises toward twenty, while implant users
plateau around seven to ten and go no further [3]. Shannon's earlier work
established the floor, intelligible speech needs surprisingly few channels in quiet
[5]. Noise is where the gap opens.
What that costs in a room
Spectral resolution is exactly the resource you spend on separating concurrent talkers. With it
degraded, implant users typically need a signal-to-noise ratio roughly 10–15 dB more
favourable than a normal-hearing listener to reach comparable sentence recognition, and
restaurants, classrooms and open offices sit well below that line.
The head-shadow effect contributes about 6.4 dB of the binaural benefit that unilateral
implant users lose outright, one reason bilateral fitting helps, and why an external array
of microphones helps further.
What we heard
The failure recipients describe is not the one the literature measures.
Everything above this line comes from published work. This section does not. It comes from
conversations with cochlear implant recipients about what actually goes wrong in a room, and it
changed what we decided to build. We are stating it plainly as what it is: a small number of
informal conversations, not a study, with no protocol and no sample size worth quoting.
Nobody said "I could not hear"
They said they nod. They follow most of a conversation, lose a sentence, and rather than stop
five people to ask, they guess and keep going. The cost is not a missed word, it is going home
unsure whether the thing they agreed to was a question, a joke or a plan.
That failure is invisible from the outside and invisible on a test. A recipient scoring well on
sentence recognition in a clinic can still be doing this at every dinner.
The second one is fatigue, not failure
The other thing we heard repeatedly is that following a room is possible and expensive. The
implant works. The person can do it. It costs so much concentration that they are wiped out
afterwards, and the rational response is to go to fewer rooms.
An implant that succeeds on every clinical measure can still lose someone their social life
through pure effort cost. That is a device outcome, and nothing in the standard battery
captures it.
Why this changed what we build
Both failures share a shape. Neither shows up in word recognition scores, which is what the field
optimises, and both are caused by the same thing: the listener is doing separation work that the
device should be doing. Effort is the product spec, not intelligibility.
It is also why we are not building captioning glasses. Reading a transcript solves comprehension
and makes effort worse, because now you are decoding text and watching a face at the same time.
If we ever measure a Leo that raises sentence scores while raising listening effort, we will
report that as a failure, and we are stating that here so it is on record before we have data.
The insight
Deaf listeners are already exceptional multisensory integrators.
This is the finding our entire product thesis rests on, and it is not ours, it is well replicated.
The visual channel in an implant user is not a fallback. It is load-bearing.
Seeing the talker is worth more the worse the noise gets
Percentage-point gain in word intelligibility from watching the speaker's face, at two
signal-to-noise ratios. The size of the closed response set changes how much vision can
contribute.
Sumby & Pollack, JASA 26(2), 1954 [4].
Spondee words, 129 listeners, each point pooled from 450 determinations. The paper reports the
−30 dB gap as ranging "from 40 percent for the 256-word vocabulary to 80 percent for the
8-word vocabulary," and states that under noise-free conditions there is little difference
between the two presentation conditions, plotted here as approximately zero.
Vision carries speech
Sumby and Pollack's 1954 measurement still sets the ceiling: at low SNRs, seeing a talker's face
is worth as much as adding roughly 15 dB of clean signal [4].
Integration is plastic
Rouger et al. found implant recipients outperform normal-hearing controls at lipreading and fuse
audiovisual speech more effectively, a durable reorganization, not a temporary strategy
[6].
Timing is the constraint
Audiovisual speech fuses inside a temporal binding window of roughly 200 ms, and asymmetrically,
listeners tolerate audio lagging vision far better than the reverse [7].
The product
Leo
AI glasses built for cochlear implant recipients. Leo does two things a sound processor cannot do
on its own: it hears the room from a wider baseline than a single behind-the-ear microphone, and
it renders what it hears into the wearer's visual field, where the evidence says it will actually
be used.
Six microphones, one baselineDK-1
A behind-the-ear processor gets one microphone and a head's worth of acoustic shadow. Leo
distributes six MEMS capsules across the temples and brow bar, giving roughly a 140 mm horizontal
baseline, enough aperture for meaningful directivity in the 500 Hz–4 kHz band where speech
information concentrates.
A minimum-variance distortionless-response beamformer steers toward the attended talker; a
second, wider-aperture estimate maintains a running map of competing sources so the system can
switch targets without re-converging from scratch.
Why not just more gain
Amplification raises target and masker together. Directional gain does not, it improves the
ratio. For a listener whose spectral resolution is capped by the physics of current spread, ratio
is the only lever that still moves the needle.
Directivity is a geometry problem before it is a software problem. Glasses happen to be
the only consumer form factor that already spans the head.
Everything stays on the frame
Separation, diarization and transcription run on an on-frame NPU. No audio leaves the device.
This is a privacy commitment, but it is a latency commitment first, a round trip to a phone,
let alone a server, spends the entire budget before any processing begins.
The stack is a causal, streaming design end to end: no model in the chain is permitted to look
ahead more than 32 ms, which rules out most published offline separation architectures and
drives essentially every engineering trade-off downstream.
The 120 ms budget
Our end-to-end design target from acoustic wavefront to implant stimulation is
under 120 ms, chosen to sit inside the audiovisual binding window with margin
[7]. Broadcast standards put the threshold of noticeable audio lag near
125 ms [8]; we treat that as a hard ceiling, not a goal.
Every stage in the pipeline below carries an allocated slice of that budget, and any model that
cannot meet its allocation does not ship, regardless of accuracy.
Compatible, not competing
Leo does not replace a sound processor and does not touch the implant's stimulation strategy.
It presents itself as an external audio source over the accessory pathways implant manufacturers
already expose, the same class of interface used by remote microphones and streaming accessories.
The processed stream is delivered as clean, pre-separated audio. The recipient's existing map,
fitted by their audiologist, does the rest. Nothing about a Leo user's clinical relationship
changes.
Vendor surface
Three manufacturers account for the overwhelming majority of implants in the field, each with its
own accessory protocol and its own telecoil and streaming behaviour. Leo abstracts these behind a
single internal coupling layer so the acoustic stack never encodes vendor specifics.
Interoperability work is ongoing and unannounced. We are not claiming certification with
any implant manufacturer, and we will not until it exists on paper.
Captions with attribution
A monocular waveguide renders transcribed speech near the lower periphery. Critically, each line
is tagged to a spatial source, the system knows which direction a phrase arrived from, so the
wearer knows who said it without breaking gaze to check.
Text is never placed over a talker's mouth. Given how much of the signal implant users extract
from articulation [6], occluding the face to display a caption would
cost more than the caption returns.
Designed against the McGurk failure mode
McGurk and MacDonald demonstrated that mismatched audio and visual speech does not degrade
gracefully, it produces a confidently perceived third syllable that was never present in either
stream [9].
A device that puts audio and text in front of the same person at slightly wrong times is not
neutral; it manufactures errors. This is the single strongest argument for the latency budget
being treated as a safety requirement rather than a performance one.
The channel that runs outwardResearch
Everything else in Leo carries the room to the wearer. This carries the wearer to the room.
Two downward-tilted global-shutter cameras under the bridge watch the signing space; an
on-device pose and hand model reads it; the result is spoken aloud.
For a deaf signer talking to someone who does not sign, that removes the phone, the notepad
and the interpreter from the loop. It is the most valuable thing we could build and the
least finished, see the full write-up for what is actually hard about it.
Translation, not transcription
American Sign Language is a language, with its own grammar, morphology and word order. It is
not English on the hands. A system that maps signs to English words in sequence produces
something between broken English and nonsense.
We are not shipping open-domain sign translation, and we would be sceptical of anyone
who says they are. The DK-1 target is fingerspelling plus a constrained phrase set.
Signal path
Wavefront to stimulation
Six stages, one budget. Allocations below are DK-1 design targets, the budget we are writing
software against. They have not been measured on hardware, because the hardware does not exist yet.
1
Capture
Six MEMS capsules sample at 48 kHz with sub-sample-synchronised clocks. Phase coherence across
the array is the whole game; a drifting clock destroys directivity faster than any amount of
self-noise.
≈ 4 ms
2
Spatial analysis
Steered-response power localisation builds a running map of active sources around the wearer,
updated every 16 ms. Head motion is compensated from the onboard IMU so the map lives in room
coordinates, not head coordinates.
≈ 16 ms
3
Separation
A causal, streaming separation network isolates the attended talker from concurrent speech and
diffuse background, the descendant of a line of work that took source separation from an offline
trick to a real-time one [15][16]. Lookahead is
capped at 32 ms, which eliminates most published offline architectures outright.
≈ 38 ms
4
Attention selection
Gaze direction, head orientation and streaming diarization, who is speaking, when, combine
into a target estimate and the speaker tag each caption carries [17].
When confidence is low the system widens rather than guesses; a wrong lock is far more
disorienting than a broad one.
≈ 12 ms
5
Delivery
Clean audio is streamed to the recipient's processor over the manufacturer's accessory
pathway. Leo hands over a better signal; the existing clinical map converts it to stimulation
exactly as it always has.
≈ 28 ms
6
Visual render
Transcription and speaker attribution reach the waveguide on a deliberately separate path,
time-aligned to the audio stream so text and sound land inside the same binding window.
≈ 18 ms
Budget check116 ms of 120 ms allocated
Four milliseconds of headroom is not comfortable, and we say so plainly. Reclaiming it is the
central engineering objective of the DK-2 revision.
Sign → speech
Giving the wearer a voice in the room
A cochlear implant solves one direction of the conversation. It does nothing for the other.
A deaf signer speaking to someone who does not sign still needs a phone, a notepad or a human
interpreter. Leo's outward channel is our attempt at removing that, cameras watch the signing
space, an on-device model reads it, and the frame speaks. In a group, that is the difference
between answering in real time and waiting for a turn that never comes.
A
Capture
Two global-shutter cameras sit under the bridge, tilted down toward the signing space in
front of the torso. Global shutter matters here: signing is fast, and a rolling shutter
skews a moving hand badly enough to change the handshape.
60 fps
B
Pose & handshape
A hand and upper-body pose model runs on the same NPU as the audio stack, producing joint
positions rather than raw video. Frames are discarded immediately, nothing that could
reconstruct a conversation is retained or transmitted.
on-device
C
Recognition
Fingerspelling and a constrained phrase set, decoded as a sequence. This is deliberately
narrow. Continuous, open-domain sign translation is not a solved problem and we are not
going to pretend otherwise.
streaming
D
Speech out
Synthesised through a small frame speaker, or to a paired phone when the wearer would
rather the voice came from the table than from their head. The wearer picks the voice, and
the wearer confirms before anything is spoken.
wearer-gated
Grammar lives on the face
Non-manual markers, brow position, head tilt, mouth morphemes, carry grammatical meaning
in ASL: negation, topic marking, and the difference between a yes/no and a wh- question. A
frame worn by the signer cannot see the signer's own face, so that information has
to be inferred from the hands alone. This is the hardest structural problem in a
first-person mount, and it is why we scoped to fingerspelling first.
Self-occlusion is constant
From a head-mounted camera looking down, hands cross, overlap and foreshorten continuously.
A third-person camera sees a signing space that a first-person camera sees edge-on. Most
published sign-recognition results come from front-facing cameras at conversational
distance, which is a materially easier problem than the one we have chosen.
The benchmarks are young
Speech recognition had decades of transcribed audio before it worked. Sign language has
orders of magnitude less data, and the largest continuous corpora are small next to what
speech models train on. Progress is real and fast, but the starting point is not
comparable.
The corpora this rests on
Dataset
Scope
Scale
Task
WWLASL
isolated
2,000 signs
Word-level ASL recognition, over 21,000 clips. The easier problem, and the one closest to fingerspelling.
HHow2Sign
continuous
11 signers
Continuous ASL over a vocabulary of more than 16,000 English words. Much harder, much closer to real conversation.
PPHOENIX-2014T
continuous
one domain
German Sign Language weather broadcasts, the field's standard translation benchmark, and narrow by construction.
WLASL: Li et al., WACV 2020 [12]. How2Sign:
Duarte et al., CVPR 2021 [13].
RWTH-PHOENIX-Weather 2014T, introduced for neural sign language translation by Camgöz et al.,
CVPR 2018 [14]. All three are third-person
recordings; none is first-person, which is the gap our own data collection has to close.
We are building this with Deaf signers, not for them, and ASL fluency is a hiring requirement
on the team that owns it. A hearing team shipping a machine that speaks on a Deaf person's
behalf, without Deaf people holding the pen, would get this wrong in ways it could not detect.
Evidence base
What the literature establishes
The figures below are drawn from published work on cochlear implant performance, not from Leo.
We publish the baseline we are trying to move before we publish anything about moving it.
No ceiling; performance limited by the masker, not the listener.
2
VNormal hearing, 20-ch vocoder
20
~20
Continues improving with channel count up to twenty.
3
VNormal hearing, 8-ch vocoder
8
8
Approximates implant performance in quiet; diverges in noise.
4
CIImplant, 22-electrode array
22
~8
Plateaus near eight regardless of electrodes active.
5
CIImplant, 12-electrode array
12
~7
Statistically indistinguishable from the 22-electrode case.
Environment
Typical SNR
CI shortfall
Intelligibility for implant users
1
QQuiet room, one talker
+20 dB
none
2
HHome, television on
+10 dB
~5 dB
3
OOpen-plan office
+5 dB
~10 dB
4
CClassroom, group work
0 dB
~12 dB
5
RRestaurant, multi-talker babble
−5 dB
~15 dB
Channel counts follow Friesen et al. [3] and
Shannon et al. [5]. Environmental SNRs are
representative ranges for the listed settings, and the intelligibility bars illustrate the
direction and magnitude reported across the speech-in-noise literature rather than any single
pooled dataset. Individual outcomes vary enormously, duration of deafness, age at implantation
and residual hearing all dominate.
Audio alone
AImplant onlyvsBImplant + vision
In multi-talker babble
COmnidirectional micvsDBeamformed array
Leo DK-1
Hardware
Developer kit configuration. Every figure below is a design target for a frame we have not built.
Nothing here has been measured, certified or clinically verified. Treat this table as the
specification we are designing to, which is the only thing it is.
Form factorTitanium frame, 52 mm lens width, prescription-compatible48 g
Microphones6 × digital MEMS, 66 dB SNR, sample-synchronised across the array140 mm baseline
ComputeOn-frame NPU; no audio or transcript leaves the device4.2 TOPS
Sign capture2 × global-shutter cameras under the bridge, tilted toward the signing space; pose extracted on-frame, frames discarded60 fps
DisplayMonocular geometric waveguide, lower-peripheral placement, never over a talker's face22° FOV
Implant couplingManufacturer accessory-streaming pathways, abstracted behind a single internal layer3 vendor targets
End-to-end latencyAcoustic wavefront to implant stimulation, inside the audiovisual binding window< 120 ms
BatteryMixed conversational use with display active roughly one third of runtime8 h
IngressSplash and dust resistant; not rated for immersionIP54
Regulatory status. Leo is a developer kit. It is not a medical device, is not
FDA cleared or approved, and is not intended to diagnose, treat, cure or prevent any condition.
Nothing here should be read as a claim of clinical benefit. Anyone with a cochlear implant should
talk to their audiologist before adding any accessory to their setup.
The market
A large market that has been counted badly.
Hearing is one of the biggest unaddressed health burdens in the world and one of the smallest
device markets relative to it. The gap between those two facts is the whole opportunity, and it
is also the reason to be careful: a market that has looked obviously large for thirty years and
stayed small is telling you something about adoption, not about need.
$980B
annual global cost of unaddressed hearing loss, across health care, lost productivity and societal cost.
Published estimates for the same market in the same year, plotted as the span between the
lowest and the highest. Each span is roughly a factor of two.
Cochlear implant estimates for 2025 collected from published report summaries by Grand View
Research, IMARC, Vantage Market Research, Fortune Business Insights, Research Nester and Mordor
Intelligence; hearing aid estimates from eight comparable reports. These are commercial research
products, not peer-reviewed work, and their methodologies are not public. Definitions differ on
whether over-the-counter devices, hearables and service revenue are counted. The spread is the
finding [25].
Implant candidates who get an implant
US utilisation among people meeting traditional audiometric candidacy criteria. Under expanded criteria it falls to about 2%.
Because the category boundary is unsettled. A hearing aid was a regulated medical device until
the US over-the-counter rule took effect in 2022, and since then the line between a hearing aid,
a hearable and a pair of earbuds with a hearing feature has been a matter of opinion. Firms that
count earbuds report figures around $15B. Firms that count only fitted devices report around $9B.
We are quoting the disagreement rather than picking the flattering end of it. A number you cannot
reproduce is not evidence, and the honest version of this slide is that the addressable device
market is somewhere in the low tens of billions and growing at high single digits.
The market we can actually reachOur estimate
Leo's first customer is an adult cochlear implant recipient in a high-income market. Cumulative
US adult implants are 118,100, and cumulative implants are not living active users, so the real
number is smaller. Call the reachable beachhead on the order of 100,000 people in the US and
perhaps three times that across comparable markets.
At a $1,500 device, a price we have not set and state here only as arithmetic you can redo
with your own assumption, that beachhead is roughly $150M in the US and $600M across
high-income markets. Small on purpose. It is where the hardest version of the problem lives.
Why now
The form factor stopped being strange
Assistive glasses have failed on stigma before. Smart glasses shipments roughly doubled year
on year through 2025, and someone else's marketing budget is normalising a camera and a
microphone array on a face. A device that looks like what everyone else is already wearing does
not carry the same cost.
Separation became streamable
Real-time speaker separation good enough for a live conversation is recent, and the models
that made it practical are small enough to run causally on a frame rather than in a datacentre
[15][16]. That is a capability change, not a packaging change.
The channel is already open
Every major implant manufacturer already ships an accessory streaming pathway. We do not have
to displace an incumbent or win a clinical channel to reach a recipient. We have to be
compatible with the one they already have.
What the size of this market does not tell you. The 16% and 12.7% figures above
are the real constraint, and they are not a distribution problem waiting for a better funnel.
Devices in this category have historically failed on comfort, stigma, reimbursement and the gap
between bench performance and a real room. We think the honest reading is that the market is large
and the conversion is hard, which is why our roadmap spends 2027 and 2028 on recipients rather
than on sales.
The field
Everyone else amplifies or transcribes.
Glasses for hearing loss are not a new idea and we are not the first people here. It is worth
being precise about what already exists, because the thing that separates Leo from all of it is
narrow and easy to state: every product below either makes the whole room louder or turns it into
text. None of them decide which voice you want and send only that one into a cochlear implant.
What exists
What it does
Who it is for
Why it does not solve this
Nuance Audio EssilorLuxottica
Open-ear directional amplification built into a normal-looking frame. FDA cleared as an over-the-counter hearing aid and on sale in the US.
Perceived mild to moderate hearing loss.
Amplification, not separation. Making a room with four talkers louder gives an implant recipient four louder talkers. It is also not a route into an implant.
AirPods Pro hearing aid feature Apple
FDA-authorised hearing aid function in a mass-market earbud, with a self-administered hearing test.
Mild to moderate loss, in-ear.
Same amplification limit, plus it occupies the ear. An implant recipient's hearing does not run through their ear canal.
Captioning glasses Xander, XRAI and others
Live speech-to-text projected on the lens. Xander runs on the device without a phone; XRAI runs off a connected phone.
Severe loss and Deaf users who read English fluently.
Text instead of sound, and it costs you the talker's face. Reading captions pulls your eyes off the mouth, and the visual speech cues you lose are worth up to 15 dB of effective SNR [4]. This is the tradeoff our whole thesis is built on.
A clip-on microphone that streams one talker directly into the implant over the accessory pathway.
Cochlear implant recipients. Widely prescribed.
This is the real incumbent and the one to beat. It works, but somebody has to physically wear it, so it solves one known talker in advance and does nothing for a table of five.
The gap we are aiming at
The remote microphone is the honest baseline, not Nuance Audio and not captions. A recipient at
a dinner table today has one option that genuinely works: hand a microphone to one person and
give up on everyone else. Leo's claim is that an array on your own face can do the choosing,
with no object to hand over and no conversation to interrupt.
That is a narrow claim, and it is falsifiable. If our separation does not beat a clipped-on
remote mic in a five-talker room, we do not have a product, and that comparison is the first
measurement we intend to publish.
What we do not claim
We are not competing with Nuance Audio or Apple. They serve mild to moderate loss, a population
perhaps a hundred times larger than ours, and they serve it well. If anything, they help us:
every pair of hearing glasses sold makes the form factor more ordinary and lowers the stigma
cost of the one we are building.
The uncomfortable version: EssilorLuxottica owns Ray-Ban, has the retail network, and could
point it at implant recipients. We do not think they will, because the population is small and
the coupling work is specialised, but a well-funded incumbent deciding otherwise is the
clearest way we lose.
Founders
Who is building this
Two people so far, paired on purpose: one who has spent years turning a condition he lives
with into working assistive technology, and one who brings the clinical depth that hearing
work demands. Between us we have the two things this problem is hard to attempt without:
direct access to recipients and clinicians, and a founder who has already shipped assistive
technology for a condition nobody was going to solve for him.
K
Kiro Moussa
Hardware & systems
Kiro studies computer science, economics and data science at MIT. He was born with congenital
nystagmus, a condition with no established cure in that form, and instead of waiting for one he
built computer-vision software that counter-oscillates on-screen text to match the motion of
the eye, holding reading still for people who live with it. It began with his own sight and his
sister's data, and earned two Silver awards at a regional science and engineering fair.
The Ear Company comes from the same instinct: not assistive technology in the abstract, only
the version that actually reaches the person who needs it.
kiro.city ↗
F
Finn
Clinical & language
Finn is an incoming computer science undergraduate and works at Stanford on clinical
speech-language pathology and NLP. His regular work with deaf children brings exactly the
clinical depth a hearing product demands, the part engineering alone cannot supply.
Kiro and Finn have known each other for two years. The split is deliberate: hardware and
systems on one side, clinical and language science on the other, with the people Leo is for
kept at the centre of both.
Practically, this is why the conversations behind the section above happened at all, and it is
the route to the academic audiology partner the 2027 studies depend on. For a company whose
hardest gate is recipient access rather than engineering, that is the part of the team that is
hard to hire for.
Roadmap
How this gets to people
We are a hardware company in a clinical adjacency, which means the honest version of our timeline
is longer than the exciting version. This is the honest one.
2026, separation running off-frameCurrent
Where we actually are. The separation and sign models run on a laptop against recorded
multi-talker audio. The DK-1 array geometry is designed and not yet built, so nothing has been
characterised in a chamber and no latency figure on this site has been measured on hardware.
2026, a working LeoNext
The whole device, assembled and running unaided: six-microphone array, separation and sign both
on-frame, own battery, coupled into an implant. This is the milestone that turns every design
target on this page into a measurement, and it is what the raise is for. Until it lands, the
honest description of Leo is software plus a drawing.
2027, Coupling interoperabilityPlanned
Bench interoperability across the three major implant accessory pathways, with a published
compatibility matrix. No performance claims until this is stable.
2027, Lab studies with recipientsPlanned
Structured speech-in-noise testing with adult implant recipients under IRB oversight, run with an
academic audiology partner. Protocol pre-registered; results published whichever way they land.
2028, Take-home cohortPlanned
Longitudinal at-home use. Bench SNR gain and real-world benefit are different quantities, and the
gap between them is where most assistive hearing technology has historically died.
Nothing on this page is a clinical result. When we have one, it will appear here
with its protocol, its sample size, its confidence intervals and its null findings attached.
References
Where the numbers come from
Published sources for every empirical figure cited above. Where we state a design target rather
than a finding, it is labelled as such in the text.
World Health Organization. World Report on Hearing. Geneva, 2021., Prevalence of hearing loss and projected growth to 2050.
National Institute on Deafness and Other Communication Disorders (NIDCD). Cochlear Implants, statistics page., Global recipient counts.
Friesen, L. M., Shannon, R. V., Başkent, D., & Wang, X. (2001). Speech recognition in noise as a function of the number of spectral channels: comparison of acoustic hearing and cochlear implants. Journal of the Acoustical Society of America, 110(2), 1150–1163.
Sumby, W. H., & Pollack, I. (1954). Visual contribution to speech intelligibility in noise. Journal of the Acoustical Society of America, 26(2), 212–215.
Shannon, R. V., Zeng, F.-G., Kamath, V., Wygonski, J., & Ekelid, M. (1995). Speech recognition with primarily temporal cues. Science, 270(5234), 303–304.
Rouger, J., Lagleyre, S., Fraysse, B., Deneve, S., Deguine, O., & Barone, P. (2007). Evidence that cochlear-implanted deaf patients are better multisensory integrators. Proceedings of the National Academy of Sciences, 104(17), 7295–7300.
van Wassenhove, V., Grant, K. W., & Poeppel, D. (2007). Temporal window of integration in auditory-visual speech perception. Neuropsychologia, 45(3), 598–607.
International Telecommunication Union. Recommendation ITU-R BT.1359-1: Relative timing of sound and vision for broadcasting., Detectability thresholds for audio lead and lag.
McGurk, H., & MacDonald, J. (1976). Hearing lips and seeing voices. Nature, 264(5588), 746–748.
Cherry, E. C. (1953). Some experiments on the recognition of speech, with one and with two ears. Journal of the Acoustical Society of America, 25(5), 975–979.
Berg, K. A., Noble, J. H., Dawant, B. M., Dwyer, R. T., Labadie, R. F., & Gifford, R. H. (2020). Speech recognition with cochlear implants as a function of the number of channels: effects of electrode placement. Journal of the Acoustical Society of America, 147(5), 3646–3656.
Li, D., Rodriguez Opazo, C., Yu, X., & Li, H. (2020). Word-level deep sign language recognition from video: a new large-scale dataset and methods comparison. IEEE Winter Conference on Applications of Computer Vision (WACV).
Duarte, A., Palaskar, S., Ventura, L., Ghadiyaram, D., DeHaan, K., Metze, F., Torres, J., & Giró-i-Nieto, X. (2021). How2Sign: a large-scale multimodal dataset for continuous American Sign Language. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
Camgöz, N. C., Hadfield, S., Koller, O., Ney, H., & Bowden, R. (2018). Neural sign language translation. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)., Introduces RWTH-PHOENIX-Weather 2014T.
Luo, Y., & Mesgarani, N. (2019). Conv-TasNet: surpassing ideal time–frequency magnitude masking for speech separation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 27(8), 1256–1266., The architecture that made single-channel voice separation practical.
Li, C., Yang, L., Wang, W., & Qian, Y. (2022). SkiM: Skipping Memory LSTM for low-latency real-time continuous speech separation. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)., Streaming, low-latency separation of the kind an on-frame device requires.
Coria, J. M., Bredin, H., Ghannay, S., & Rosset, S. (2021). Overlap-aware low-latency online speaker diarization based on end-to-end local segmentation. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)., Deciding who is speaking, incrementally, as audio arrives.
World Health Organization. Global costs of unaddressed hearing loss and cost-effectiveness of interventions. Geneva, 2017., Source of the US$ 980 billion annual cost figure reported in the World Report on Hearing.
National Institute on Deafness and Other Communication Disorders (NIDCD). Quick Statistics About Hearing, Balance, & Dizziness., Hearing aid use among US adults who could benefit.
Nassiri, A. M., Sorkin, D. L., & Carlson, M. L. (2022). Current estimates of cochlear implant utilization in the United States. Otology & Neurotology, 43(5), e558–e562.
Counterpoint Research. Global smart glasses shipments, H1 and H2 2025 market briefs., Shipment growth and vendor share for the smart glasses category.
EssilorLuxottica. Nuance Audio Glasses., FDA 510(k) clearance and CE marking announced February 2025 as an over-the-counter hearing aid; US availability from April 2025. Indicated for perceived mild to moderate hearing loss.
Apple. Hearing Aid Feature, AirPods Pro., FDA-authorised over-the-counter hearing aid software feature, announced September 2024, for perceived mild to moderate hearing loss.
Xander Glasses; XRAI Glass. Live captioning eyewear., Xander performs speech-to-text on the device without a paired phone; XRAI runs captioning software on connected AR glasses via a phone. Cited as the captioning category, not as endorsement.
Commercial market research summaries, 2025–2026 editions: Grand View Research, IMARC Group, Vantage Market Research, Fortune Business Insights, Research Nester, Mordor Intelligence, MarketsandMarkets, Custom Market Insights., Cited only to show the spread between them. These are paid research products whose methodologies are not published, and we do not treat any single figure as evidence.
Mohammadi, S.-F., et al. Assessment of refractive errors, amblyopia, strabismus, and low vision among hearing-impaired and deaf students in Kermanshah., Reports vision problems in 17–30% of normal-hearing people against 44–65% of deaf and hard-of-hearing people, and states that the retina and cochlea originate from a single layer in the sixth and seventh weeks of the embryonic period. The underlying prevalence studies weight toward children and adolescents.
Usher syndrome literature. Prevalence of 4–17 per 100,000, and status as the most common inherited cause of deafblindness; cochlear implantation outcomes in Usher recipients., Reviews at PMC7502997 and PMC11487040.
Dual sensory impairment prevalence. Combined hearing and vision loss rising from roughly 1.5% at ages 65–74 to 2.6% at 75–84 and 10.8% past 85, with associated increases in dementia risk., Population-based studies including PLOS One (2013) and Fuller-Thomson et al. (2022).
Company
The Ear Company, LLC
We build AI hardware for bionic enhancement. We started with hearing because it is the one sense
where a widely deployed neural implant already exists, has been refined over four decades, and is
still bottlenecked by something a well-designed external device can genuinely help with.
Compatible by default
We build things that attach to the clinical world rather than replacing it. No recipient should
have to abandon a working map, a trusted audiologist, or a device they have spent years learning.
On-device or not at all
Conversation is the most sensitive data a person produces. Leo processes it on the frame and
keeps it there. That constraint shapes the architecture rather than being bolted onto it.
Publish the baseline first
We put the literature we are trying to beat on our own website before we put up a single number
of our own. If our results do not clear that bar, that will be visible here too.
Where we are
The Ear Company, LLC is a limited liability company headquartered in San Francisco, California.
Hardware, acoustics and machine learning are in one room, which for a latency-bound product is
less a culture choice than an engineering requirement.
Be first in line
Get Leo before anyone else.
Join the waitlist for early access to the DK-1 developer kit and the occasional, honest
update on where the engineering actually stands. No hype, no spam, unsubscribe in one click.
If you invest: we are raising a pre-seed to build a fully working Leo, the whole
device on a face rather than a bench, and then measure it in a real room against a clipped-on
remote microphone, the incumbent we actually have to beat. That comparison is the next thing that
will appear on this page, whichever way it lands. We are happy to walk through the off-frame demo
and the latency budget on a call before you decide anything.