20 examples where an extra cue can help
Typing errors can leave a word garbled or turn it into another valid word. A brief whisper or silent lip movement can add evidence to help recover what you intended. See what the cue contributes in selected examples below.
U I O sit side by side on a QWERTY keyboard. That makes pairs such as love/live, this/thus, if/of, and pick/puck easy to mistype. These tiny slips can be frustrating because a real word may pass a spelling check, and context does not always rescue the mistake. Faster or less precise typing can also leave a word garbled, making the intended spelling harder to recover from the typed letters alone.
10 isolated word examples
| # | You meant | Typed by mistake |
| 01 | love | live |
| 02 | this | thus |
| 03 | if | of |
| 04 | shirt | short |
| 05 | pick | puck |
| 06 | restaurant | restrant |
| 07 | avocado | abcado |
| 08 | appointment | apointmrnt |
| 09 | prescription | prescriotn |
| 10 | accommodation | acxomodatn |
10 short sentence examples
| # | You meant | Typed by mistake |
| 11 | I love my life here. | I live my life here. |
| 12 | I left it in the car. | I left it on the car. |
| 13 | I like the shirt dress. | I like the short dress. |
| 14 | We discussed the renovation. | We discussed the renovstn. |
| 15 | Please bring the coat. | Please bring the boat. |
| 16 | The team is ready. | The tea is ready. |
| 17 | Please send the confirmation. | Please send the confrmatn. |
| 18 | I’ll go with the latte. | I’ll go with the latter. |
| 19 | It’s a casual relationship. | It’s a causal relationship. |
| 20 | What about Wednesday? | What about wrfnesdy? |
In the valid-word examples, both versions are grammatical but mean different things. A language model can favor one, yet still miss your intent.
What the cue contributes
A cue need not reveal the whole word to help resolve it. Even faint whispers often carry clues to syllable count, rhythm, critical vowel or consonant sounds (phonemes), and timing. An optional, brief camera capture can add visemes, visible lip shapes associated with speech sounds, such as closure or rounding, and the movements between them. LipConfirm combines the available cues with the typing to help infer the intended word; audio and lip evidence can also reinforce one another.
Whisper · A distinguishing vowel
love / live, this / thus
In “I love my life here,” a faint “uh” vowel cue can favor love over live, pronounced “liv.” The same vowel contrast helps separate thus from this. Within each pair, the consonants and syllable count match. The typing narrows the choice; the vowel adds evidence context may lack.
Whisper · Recovering garbled typing
avocado / abcado
If you type abcado but whisper avocado, the surviving letters still provide part of the word’s structure. The whisper can add vowel sounds and syllable rhythm that the garbled spelling does not preserve. LipConfirm can evaluate both together, including in a multimodal model that infers the word directly from the combined inputs. The whisper need not identify the word on its own.
Silent lips · An ending the text missed
team / tea
If you type “The tea is ready” but articulate team, the lips close for the final “m.” A captured closure timed to that word ending can support team over tea. LipConfirm can weigh that visible cue alongside the shared typed letters; a resting mouth closure alone would not establish the word.
Whisper or lips · A different beginning
coat / boat
A captured initial “k” sound can favor coat over boat. In the other direction, boat begins with a “b” made by closing and releasing both lips, while “k” is formed farther back in the mouth. A captured lip closure and release at the word’s start can support boat when evaluated together with the typing.
Whisper · Syllable pattern and timing
casual / causal
When casual is articulated with three syllables, its rhythm differs from the usual two syllables of causal. The vowel and middle consonant sounds also differ. LipConfirm can weigh those clues against the nearly identical typed letters. Pronunciation varies, so syllable count is supporting evidence rather than a fixed rule.
Whisper or silent lips · A distinguishing vowel
if / of
These words differ by the neighboring “i” and “o” keys. Their vowel visemes can also differ: the “ih” in if can have a different visible shape from the more open vowel in a fully articulated of. LipConfirm can weigh that shape and its transition into the final consonant against the typing, with whisper audio when available. The visible contrast depends on pronunciation and articulation.
These walkthroughs illustrate how cues can support correction. Pronunciation and capture quality determine which evidence is available. When the combined evidence is insufficient, LipConfirm can retain the typed word or offer alternatives.
How it works
LipConfirm is designed to bring typed input and supporting whisper or lip evidence into existing keyboard, autocorrect, spell-checking, and language-model infrastructure. Technical demonstrations and IP discussions available. Desktop prototypes cover both whisper-assisted and lip-assisted correction.
Capture a brief cue
Whisper or silently mouth a word as you type it, or after you notice a mistake. LipConfirm captures the whisper through the microphone, or briefly uses the camera to track silent lip movements when you request help. The whisper does not need to produce a reliable dictation transcript, and the camera need not stay on throughout typing.
Two ways to combine the inputs
Rerank candidates: Start with suggestions generated from typing, then use whisper or lip evidence to rescore or extend those candidates.
Infer with a multimodal model: Evaluate typing and whisper or lip evidence together in a learned model to infer the intended word. The word need not already appear in a separate autocorrect list.
Suggest, confirm, or replace
The combined evidence can update a suggestion, confirm a word, or correct a valid word that was typed unintentionally. Confidence controls matter for familiar words as well as garbled input: when the evidence is insufficient, the system can retain the text or offer alternatives.
How faint whisper cues reach correction
A faint whisper often retains useful clues to syllable count and rhythm, critical vowel or consonant sounds (phonemes), and sound timing. LipConfirm evaluates the available acoustic evidence together with the original typing, including garbled text. The whisper need not identify the word independently: a partial cue can help resolve the intended word when combined with the typing. A desktop demo of whisper-assisted correction is available.
The capture path must preserve that weak signal. Depending on the system, speech-detection gates or audio filtering can exclude faint utterances before they reach recognition. A user-enabled typing session or requested capture can provide a path for evaluating the retained acoustic evidence with the text. A blank dictation result alone does not identify which stage failed.
Sources: whispered speech
Phonetic background: research on whispered speech documents vowel resonances and sound durations despite the absence of normal vocal-fold vibration. This helps explain why incomplete acoustic evidence can remain informative; it does not establish a minimum usable volume for LipConfirm.
How visemes help resolve typed words
Whisper remains the primary auxiliary cue. If you prefer not to whisper at all, LipConfirm can briefly use the camera to track your silent lip movements when you request assistance. It is designed to recognize visemes, the transitions between them, and visual syllable patterns such as the rhythm of mouth opening and closing. These provide alternative visual evidence to evaluate alongside the original typing.
The table below shows examples of visemes detected by the current LipConfirm desktop demo.
| Viseme category | Visible pattern and example sounds |
| CLOSED Bilabial | Both lips meet, as for /p/ in pat, /b/ in bat, and /m/ in mat. |
| LABIODENTAL | The lower lip contacts the upper teeth, as for /f/ in fan and /v/ in van. |
| SPREAD | The lips spread for an “ee” sound, /iː/, as in see. |
| OPEN | The mouth opens for an “ah” sound, /ɑ/, as in father. |
| ROUNDED | The lips round or protrude for “oo,” /uː/, as in food, or “oh,” /oʊ/, as in go. |
| MID | A relaxed, moderately open mouth shape for “ih,” /ɪ/, as in sit or if. |
LipConfirm can combine these shapes, their transitions, and visual syllable patterns with the typed letters. For example, a final CLOSED shape can support team over tea; a SPREAD or MID vowel shape can help separate words with “ee” or “ih” sounds.
Sources: visual speech and articulation
These references provide background for the visual speech cues and articulation described above. LipConfirm’s preliminary correction results are described separately in the accuracy FAQ.
What the demos show
- Garbled input and everyday mix-ups. The interactive demo includes both malformed spellings and valid-word errors such as love/live. The videos show selected lip-assisted correction scenarios.
- Typing stays primary. Lip cues in the videos guide a typed correction. Whisper assistance applies the same principle using quiet audio.
- No need to repeat everything. An optional cue can accompany any word you want help with, including a familiar word typed incorrectly. You can also add the cue after you notice a mistake. You do not need to whisper or mouth the rest of the message.
- Note: The camera panel is a behind-the-scenes visual aid. An integrated keyboard need not display the feed, and requested visual correction need not keep the camera running throughout typing.
What the demos show
- Garbled input and everyday mix-ups. The interactive demo includes both malformed spellings and valid-word errors such as love/live. The videos show selected lip-assisted correction scenarios.
- Typing stays primary. Lip cues in the videos guide a typed correction. Whisper assistance applies the same principle using quiet audio.
- No need to repeat everything. An optional cue can accompany any word you want help with, including a familiar word typed incorrectly. You can also add the cue after you notice a mistake. You do not need to whisper or mouth the rest of the message.
- Note: The camera panel is a behind-the-scenes visual aid. An integrated keyboard need not display the feed, and requested visual correction need not keep the camera running throughout typing.
FAQ
Why not just whisper during dictation as it improves?
You may prefer typing to keep most of a message unspoken or avoid disturbing others. The whisper you are comfortable making at normal typing distance may also be too faint for reliable dictation. LipConfirm adds the evidence in your actual typed input, so the same uncertain whisper can help resolve a word without having to identify it independently. This can be useful even as dictation improves.
Can typing plus a faint whisper be more accurate than that whisper alone?
LipConfirm is showing that combining faint whisper cues with typed input can help identify the intended word, including when the typing is garbled. These are early findings from selected desktop demonstrations. A representative accuracy percentage across users, devices, noise levels, and normal phone typing distances has not yet been established. Broader evaluation needs to count both recovered words and unwanted corrections.
Does this apply only to hard-to-spell words?
No. Familiar words can be mistyped into other valid words, including love/live and this/thus. “I love my life here” and “I live my life here” are both grammatical, so even a strong language model may not settle the user’s intent from context alone. A whisper or lip cue can supply additional evidence; confidence checks determine whether to change the word.
What if dictation does not recognize my faint whisper?
The microphone can capture a faint whisper even when dictation cannot reliably transcribe it. Useful acoustic clues can remain in that audio even if dictation returns no text or the wrong word. LipConfirm uses the captured signal together with your original typing to help infer the intended word, without requiring a reliable standalone transcript. Try the progressively fainter whisper check at your usual typing distance.
Does whisper assistance keep a message private?
A very faint whisper can help protect privacy by making a correction harder to overhear. Whispers vary in volume, and a faint cue can be much less audible nearby than an ordinary whisper. LipConfirm combines the captured signal with your typing, so even a cue too faint for reliable dictation can help clarify the intended word. Most of your message stays unspoken, and optional silent lip assistance provides a silent path.
For data privacy, the architecture supports local processing and discarding raw audio or mouth-region frames. The host platform determines the implemented data handling.
What is whisper-level audio?
Here, whisper-level audio means quiet acoustic evidence, including faint, low-amplitude or breathy speech. The aim is to use a comfortable brief cue while keeping the phone where you normally type. LipConfirm evaluates that evidence with the typed input; a complete, correct standalone transcript is not required. Usable volume depends on capture settings, microphone distance, background noise, and the model.
Does LipConfirm require a camera?
No. Whisper-assisted correction combines microphone input with typing and does not need a camera. Optional lip assistance can provide a silent cue or reinforce a quiet utterance.
For silent lip cues, does the camera run throughout typing?
Continuous camera use is not required. In the on-demand approach, the camera stays off until lip assistance is requested, captures a short clip of the word, then stops. Longer visual sessions remain an optional, user-enabled configuration.
Can the microphone stay available while I type?
Yes, in an opted-in typing session. The microphone can remain available so an occasional quiet cue accompanies a word as you type it, including a familiar word that is easy to mistype. Alternatively, it can activate only when assistance is requested. There is no requirement to whisper every word.
Can I correct a word after it appears?
The architecture supports associating a new cue with the word just typed and reconsidering the correction using the retained typing evidence. Optional lip assistance can use a brief camera capture requested for that word.
Can whisper and lip cues work together?
Yes. One quiet articulation can supply audio and lip cues. The architecture combines both with typed evidence and adjusts their influence according to reliability. An unavailable or unreliable channel can contribute less or be omitted.
What is demonstrated today?
Desktop demonstrations of whisper-assisted and lip-assisted correction are available. The interactive website demo is an illustrative simulation with preset examples and scores; it does not access your microphone or camera or measure live recognition accuracy.
Does this require the internet or cloud processing?
LipConfirm is designed to support on-device correction without sending raw mouth-region frames or audio to the cloud. Raw sensor data can be discarded after feature extraction or scoring. Deployment choices depend on the host platform and model.
Will this drain battery or slow down typing?
The design supports brief camera use and session-based or requested microphone capture. Actual battery use and latency depend on sensor settings, recognition models, and the device.
How does the extra evidence enter autocorrect?
LipConfirm can rescore or extend suggestions generated from typing, or infer a correction through a learned multimodal model that evaluates the inputs together. A proxy-string pathway can turn speech-related clues into spelling-like inputs for an existing correction engine. The interactive demo illustrates late and early fusion as two distinct strategies.
Do I need to perfectly mouth the words?
Text-assisted correction can use partial cues, such as syllable patterns and mouth shapes, rather than independently recognizing every word from lips alone. Useful visual evidence still depends on articulation and capture quality. Optional personalization can adapt scoring to a user.
Is LipConfirm only for mobile phones?
The architecture supports keyboard-based correction on phones, tablets, desktop systems, wearables, and extended-reality devices. Each platform can use the available, user-permitted combination of typed, acoustic, and visual evidence.
Will I still see suggestions?
A high-confidence result can be inserted or confirmed automatically. When the evidence is uncertain, the system can offer alternatives, retain the current text, or wait for another cue. The user does not need to see the internal candidate scores or camera feed.
Is this patented?
All 30 claims of the U.S. parent application have been allowed, and U.S. Patent No. 12,737,540 is scheduled to issue September 15, 2026. Two U.S. continuation applications and a PCT international application are pending. One continuation has been granted Track One prioritized examination. See Patent & IP for details.