Speak 2 Move A Vision-Language-Based Semantic Mapping Framework For Autonomous Navigation In Assistive Robots

Traditional control interfaces for powered assistive devices are often inaccessible to individuals with severe mobility impairments. Embodied AI, which integrates perception, reasoning, and action within robotic systems, offers a promising approach to improve accessibility and autonomy. This work presents Speak2Move, an embodied AI framework that enables intuitive voice-based control of assistive mobility devices. By integrating a large language model with object detection, semantic segmentation, SLAM, and autonomous navigation, Speak2Move creates a closed-loop pipeline connecting perception, reasoning, and action. The system incorporates a novel semantic mapping method and is deployed on a physical robot equipped with an RGB-D camera. Performance is evaluated through nine real-world navigation trials assessing spatial reasoning, obstacle avoidance, semantic understanding, and high-level task planning. A user study with five participants showed that Speak2Move consistently reduced cognitive workload during navigation, highlighting the potential of embodied AI to enhance accessibility, independence, and user experience in assistive mobility.