Vision Scribe Research Sign in

Reading speech from silent video.

We build the consented video data that visual speech recognition needs, and we train models on it from scratch. Every clip is licensed, every contributor is a volunteer, and every model is clean room.

Get in touch Contribute recordings

What we are building

A licensed dataset

Video recorded by volunteers who agreed, in writing, to how it would be used. No scraped faces and no ambiguous provenance.

Models trained from scratch

Nothing pretrained and nothing borrowed from research code that forbids commercial use. What we ship, we can account for.

Built to run inside products

The models are made to sit behind other software, from captioning and conferencing to accessibility tools. Partners license the data, the model, or both.

Where it is used

Speech that audio cannot carry

People who have lost their voice to illness or surgery still form words. A camera can read them when a microphone cannot.

Rooms that are too loud, and too quiet

A factory floor defeats a microphone. So does an open office where speaking aloud is not an option.

Captions when the audio fails

Dropped packets, a muted mic, a bad connection. The picture keeps carrying the sentence after the sound stops.

Privacy is the design, not the policy page

  • Your face is cropped in your own browser. Only the lower face leaves the page, and the rest of the picture is never uploaded.
  • Your microphone confirms you read the sentence on screen. The model never receives audio as an input.
  • Consent is versioned and timestamped. When the terms change you are asked again, and the record shows which text you agreed to.
  • Deletion is yours to run. It removes your clips, your profile, and the cropped copies derived from them.
  • Our analytics are cookieless. We count visits to the site, and nothing that identifies you or follows you anywhere else.

Read the full consent terms