Home Artificial Intelligence Google DeepMind Brings Sign Language Translation to Phones With SL2T – Unite.AI

Google DeepMind Brings Sign Language Translation to Phones With SL2T – Unite.AI

by admin
Google DeepMind Brings Sign Language Translation to Phones With SL2T – Unite.AI

Google DeepMind has shipped a sign-language-to-text model, SL2T, inside two consumer Android products, the first time sign language AI has reached a shipping phone feature. The model powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11 phones, launching with American Sign Language to English translation on August 12, 2026.

The feature works anywhere a user would normally type. A Deaf signer can dictate a web search, draft a message or document, or query Gemini by signing to the phone’s camera instead of typing. In Live Transcribe, signers can respond in signed conversation rather than typing back and forth. Google says testers found signing in ASL faster and more natural than typing in English.

SL2T is the first model Google describes as a breakthrough in quality and generality for sign language translation. It was trained on more than 100,000 hours of signing data spanning over 50 sign languages, with roughly a quarter of the data in ASL. Training jointly across languages, dialects, and proficiency levels produced a model that outperforms single-language systems in the company’s experiments, according to the announcement.

What the Model Sees and How It Translates

The system does not process raw video of the signer. An on-device model, MediaPipe Holistic, tracks 130 key points on the face, body, and hands frame by frame, and only those geometric coordinates leave the device for translation. The original camera feed is discarded immediately, a design Google frames as a privacy safeguard.

From that coordinate stream, SL2T generates English text directly, skipping the intermediate annotation layer known as glosses that most prior sign language translation systems relied on. Glosses assign fixed labels to individual signs, which limits vocabulary and loses the non-manual markers and spatial constructions that carry grammatical meaning in sign languages. Removing that bottleneck lets translation quality scale with training data rather than with a hand-built label set.

Sign languages are independent natural languages with their own grammars, not signed versions of spoken English. That makes the task machine translation between two languages, plus a hard computer vision problem: the model has to track simultaneous movement of the hands, arms, torso, head, and face at high frame rates. Earlier attempts at sign language technology, such as sensor gloves, could not address either requirement.

Benchmark Results and Remaining Error Patterns

On the FLEURS-ASL benchmark, which measures ASL-to-English translation of complex, abstract language recorded in studio conditions by Certified Deaf Interpreters, SL2T scores 70 BLEURT zero-shot, well above previously reported results, according to Google. The company also reports 74 BLEURT on a one-handed signing variant of the benchmark and 85 BLEURT on PRESTO-ASL, a dataset of assistant-style phrases signed to a docked tablet. On FSboard, a fingerspelling dataset collected from Deaf signers on mobile devices, the model reaches 64 percent exact-match accuracy. These are vendor-reported figures; no independent evaluation has been published.

Google published side-by-side examples from FLEURS-ASL showing fluent output alongside identifiable failure modes. The model occasionally substitutes rare signs and rapid fingerspelling, drops passive constructions and classifier depictions, and loses tense when context is missing. One example translates “This fully feathered, warm blooded bird of prey” as “This creature is warm-blooded, eats grey,” a fingerspelling error substituting “grey” for “prey.”

The model also shows residual hallucination behavior, generating text when no one is signing, such as when a second person enters the frame or the signer pauses. Google says the release version significantly reduced these errors compared with the pre-release builds that Deaf evaluators tested in July 2026, but has not eliminated them.

Where Google Says the Tool Must Not Be Used

The release carries an unusual layer of governance. Google DeepMind convened an advisory body, the AI Sign Language Advisory Committee, bringing together Deaf organizations including the National Association of the Deaf, the World Federation of the Deaf, the Deaf Professional Arts Network, and Rochester Institute of Technology’s National Technical Institute for the Deaf. The committee co-authored a joint impact report that spells out what the technology is for and where it stops.

The report defines SL2T 1.0 as an assistive draft-generation tool for low-stakes, informal settings: messaging, search, navigation, note-taking, and casual one-on-one conversations. It explicitly rules out medical consultations, legal proceedings, police interactions, classroom instruction, job interviews, and government benefits determinations. It also states that the tool does not satisfy legal obligations for reasonable accommodations under the Americans with Disabilities Act or equivalent frameworks, and that it must not be used by institutions to substitute for certified human interpreters.

The signer stays in the loop throughout. Both apps show a text preview the user can review and edit before sending, and in Live Transcribe the Deaf user controls when translated text is shown to a conversation partner. The report notes this safeguard assumes functional English literacy to catch errors.

How SL2T Performs in the Model’s Own Limits Document

The impact report catalogs the model’s technical boundaries in detail. Accuracy degrades in low light, harsh backlighting, or when clothing reduces contrast against the hands and face. The pose tracker follows only the largest subject in frame and cannot handle multiple signers. Performance drops at extreme camera angles or when the signer lies sideways. Input clips are capped at 60 seconds, and each clip is translated independently, with no memory of earlier conversation turns. The model was not trained or evaluated on signers under 18. It supports no custom word lists, no name signs, and no dialect preferences.

Near-term updates already on the roadmap include dictation of punctuation, newlines, emojis, and basic editing commands, plus conversation memory across sequential clips in Live Transcribe. Longer-term research targets include better sensitivity to non-manual markers such as facial expressions and eyebrow position, multi-signer support, and expanded data collection across regional ASL dialects and Black ASL.

The project originated with Sam Sepah, a Deaf Googler, and Deaf perspectives were built into data collection, user studies, and evaluation. Google describes SL2T’s launch as the beginning of a longer effort: additional sign languages, sign language generation, and broader frontier capabilities are all listed as active work. The feature ships on Pixel 11 at no additional cost, with more devices to follow.

Source Link

Related Posts

Leave a Comment