cd /news/artificial-intelligence/discosign-discourse-aware-text-to-si… · home topics artificial-intelligence article
[ARTICLE · art-126905] src=machinelearning.apple.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

Researchers Vasileios Baltatzis, Mert Inan, Connor Gillis, Raja Kushalnagar, Lorna Quandt, Leah Findlater, and Colin Lea introduced DiscoSign, a modular Large Language Model-based framework for discourse-aware text to sign language gloss translation, published in September 2026. DiscoSign addresses three phenomena — spatial coreference resolution, Question-Answer Clauses (QACs), and concept-gloss consistency between English concepts and American Sign Language (ASL) signs — and the authors report that discourse-aware processing significantly improves spatial consistency and entity tracking over sentence-only translation while maintaining competitive single-sentence gloss translation quality. The work also introduces novel evaluation metrics for discourse-level quality and is described as the first systematic framework for discourse-level text to sign language gloss translation with corresponding evaluation methodology.

read2 min views1 publishedSep 11, 2026
DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
Image: Apple ML Research

content type paperpublished September 2026 DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

AuthorsVasileios Baltatzis‡, Mert Inan‡†, Connor Gillis, Raja Kushalnagar§, Lorna Quandt§**, Leah Findlater, Colin Lea

Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for discourse-aware text to sign language gloss translation grounded in linguistic research. We address three key phenomena within our modular Large Language Model (LLM)-based translation framework: (i) spatial coreference resolution, where entities maintain consistent spatial locations throughout discourse; (ii) Question-Answer Clauses (QACs), pseudocleft structures serving specific discourse functions; and (iii) concept-gloss consistency, ensuring stable mappings between English concepts and American Sign Language (ASL) signs. Traditional translation metrics fail to capture discourse-level quality, so we introduce a suite of novel evaluation metrics designed to assess each dimension of discourse coherence addressed by our framework. Experiments on sentence-level and discourse-level datasets show that our approach for discourse-aware processing significantly improves spatial consistency and entity tracking relative to sentence-only translation, while maintaining competitive single-sentence gloss translation quality. Our work establishes the first systematic framework for discourse-level text to sign language gloss translation with corresponding evaluation methodology.

Bootstrapping Sign Language Annotations with Sign Language Models

April 30, 2026research area Accessibility, research area Computer Visionconference CVPR AI-driven sign language interpretation is limited by a lack of high-quality annotated data. New datasets including ASL STEM Wiki and FLEURS-ASL contain professional interpreters and 100s of hours of data but remain only partially annotated and thus underutilized, in part due to the prohibitive costs of annotating at this scale. In this work, we develop a pseudo-annotation pipeline that takes signed video and English as input and outputs a ranked…

Towards AI-Driven Sign Language Generation with Non-Manual Markers

March 7, 2025research area Accessibility, research area Human-Computer Interactionconference CHI

Sign languages are essential for the Deaf and Hard-of-Hearing (DHH) community. Sign language generation systems have the potential to support communication by translating from written languages, such as English, into signed videos. However, current systems often fail to meet user needs due to poor translation of grammatical structures, the absence of facial cues and body language, and insufficient visual and motion fidelity. We address these…

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @discosign 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/discosign-discourse-…] indexed:0 read:2min 2026-09-11 ·