👄 Visual Speech Recognition
Upload a silent video. Mouth movement is transcribed to text — no audio is read.
Checking connection to local API on port 6040...
⚠️ Accuracy expectations — please read
This app uses open-source pretrained models
(Auto-AVSR,
optionally AV-HuBERT).
Even the best published result is around 20% word error rate on clean,
well-lit, front-facing video — meaning roughly 1 in 5 words can be wrong.
Poor lighting, side angles, facial hair, or fast speech make this worse.
English only. See docs/ACCURACY_NOTES.md for details.