Multimodal Data for Humanoid Robots

Multimodal Data for Humanoid Robots: Vision, Language, Action, Telemetry, and Context

Ask a humanoid robot to “pick up the red mug on the left and place it in the sink,” and a remarkable amount has to happen at once. The robot must see the mug, parse the instruction, plan a motion, feel the grip pressure, and understand that “the sink” is the wet basin three feet […]