In early 2024 our transcripts came from a server in Frankfurt. They were accurate, they arrived in about twenty seconds, and every word you said passed through a machine we rented. For a voice-notes app, that last part never sat right.
What changed on the phone
Two things happened at once. Phones got neural engines fast enough to run a speech model in real time, and open speech models got small enough to fit in 180 MB without losing much accuracy. By the end of 2025 an iPhone 13 could transcribe an hour of audio in under seven minutes.
The best privacy policy is not having the data. On-device made that possible for us.
What it cost
Moving meant rewriting the transcription pipeline, retraining for accents we had only ever handled in the cloud, and accepting a slower first week while we tuned it. Our word error rate went up by 0.6 points in beta. It is now lower than the cloud version ever was.
The numbers today
Word error rate: 4.1% on our 400-note test set.
One hour of audio: under four minutes on an iPhone 12.
Data sent to Earshot servers during transcription: zero bytes.
If you turn on sync, notes travel end-to-end encrypted with a key only your devices hold. If you do not, they never leave your phone at all.
