cd /news/generative-ai/gazedit-gaze-accurate-diffusion-imag… · home topics generative-ai article
[ARTICLE · art-132238] src=arxiv.org ↗ pub= topic=generative-ai verified=true sentiment=↑ positive

GazeDiT: Gaze-Accurate Diffusion Image Generation for Eye Tracking via Spatial Conditioning

GazeDiT, a diffusion model that generates eye-tracking images for a requested 4D binocular gaze via an internally constructed spatial condition, achieves substantially lower tail gaze-label error than other diffusion baselines, approaching the error of the same frozen gaze estimator on real images. The model uses a frozen SegFormer to extract pupil and iris geometry during training and a physical eye renderer at inference to sample gaze-consistent geometries without a source image. Its generated data improved a downstream eye tracker, reducing gaze error on difficult cases from 3.05 degrees to 2.80 degrees in the smallest cohort.

by read1 min views1 publishedSep 17, 2026

arXiv:2609.17814v1 Announce Type: new Abstract: Diffusion models are increasingly used to generate synthetic training data, but precise label control remains difficult when the conditioning signal is low-dimensional and coarse. Text-conditioned images are judged by broad prompt consistency, whereas supervised training requires precise correspondence between each image and its numerical label. This is challenging in eye tracking, where a 4D binocular gaze is expressed through subtle, spatially localized pupil and iris geometry. We introduce GazeDiT, a diffusion model that generates images for a requested 4D gaze through an internally constructed spatial condition that grounds the global gaze label in this local geometry. During training, a frozen SegFormer extracts pupil/iris geometry from diverse real images, allowing the model to learn realistic appearance conditioned on that geometry. At inference, a physical eye renderer samples gaze-consistent geometries by varying anatomy and camera state, enabling diverse synthesis without a source image. GazeDiT achieves substantially lower tail gaze-label error than other diffusion baselines, approaching the error of the same frozen gaze estimator on real images. Its generated data also improves the downstream eye tracker, reducing gaze error on difficult cases from 3.05{\deg} to 2.80{\deg} in the smallest cohort.

── more in #generative-ai 4 stories · sorted by recency
── more on @gazedit 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gazedit-gaze-accurat…] indexed:0 read:1min 2026-09-17 ·