iPhone 12 and 13 now use a smaller transcription model that we trained on the same data as the full one.
One hour of audio: 3 min 52 s, down from 6 min 15 s.
Word error rate on our 400-note test set: unchanged at 4.1%.
Battery use during transcription: down 22%.