I've been looking into modern robot learning datasets, and it seems like a lot of information can now be estimated from RGB video using existing models, such as: 3D hand pose Camera trajectory Depth Segmentation Point clouds Natural language task annotations That made me wonder where the limits are. What signals still cannot be recovered reliably from video and therefore need dedicated sensors during data collection? For example: Tactile/contact sensing? Force/torque? Eye gaze? EMG? Something else? I'm interested in understanding what data collection bottlenecks still exist for manipulation and embodied AI. submitted by /u/Tricky-Promotion6784 [link] [Kommentare]
Log in Log in to comment.
No comments yet.
Comments