Scene2Intent: A Proactive Interaction Framework for Blind and Low Vision Navigation from Egocentric Video Analysis
How can we build a proactive navigation AI framework that anticipates the implicit mobility needs of BLV travelers?

Abstract
Current AI assistants for Blind and Low Vision (BLV) travelers primarily respond to explicit requests, despite many mobility needs emerging before users are able to formulate a query. We present Scene2Intent, a proactive human-centered AI framework for navigation interfaces that determines when and what assistance should be provided based on users’ mobility context. Egocentric travel videos are analyzed using multimodal AI to generate Scene–Behavior–Intent representations derived from theory-informed coding and interviews, with the resulting patterns validated by BLV participants. These patterns are translated into a three-layer interaction framework for anticipating implicit needs through continuous assistance, context-triggered guidance, and task-specific interactions.
Current Outcomes
Research, translation, and recognition.
Research Outputs
- (SCI/SSCI-1) Chen Z., Lei C., Ye Y., et al. (under review). Scene2Intent: A Proactive Interaction Framework for Blind and Low Vision Navigation from Egocentric Video Analysis. International Journal of Human-Computer Interaction (IJHCI). Submitted in August 2026.
- (CCF-A) Lei C., Wang Y., Chen Z., et al. (manuscript). SceneCue: Confirmed Proactive Task Switching for Blind and Low Vision Pedestrian Navigation. Preparing for submission to CHI 2027.
Research Translation
- Travel Assistant for the Visually Impaired ‘AccGO’: three MVP iterations and four rounds of BLV user studies.
- International College Students Innovation Competition, 2026 — participating.
- Finalist, 2026 Alibaba AI for Good Innovation Competition — participating.
Project Media
Prototypes in motion.
Travel Assistant for the Visually Impaired ‘AccGO’
An MVP demonstration developed through iterative BLV user studies.