👄 Visual Speech Recognition

Upload a silent video. Mouth movement is transcribed to text — no audio is read.

Checking connection to local API on port 6040...
⚠️ Accuracy expectations — please read

This app uses open-source pretrained models (Auto-AVSR, optionally AV-HuBERT). Even the best published result is around 20% word error rate on clean, well-lit, front-facing video — meaning roughly 1 in 5 words can be wrong. Poor lighting, side angles, facial hair, or fast speech make this worse. English only. See docs/ACCURACY_NOTES.md for details.