研究突破 arXiv cs.AI
用视觉提示让机器人学会便利店抓取摆放 Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT
精读摘要
便利店的机器人抓取摆放任务充满挑战:物品密集、相互遮挡、颜色形状尺寸纹理各异,都让轨迹规划与抓取更困难。研究提出一种感知-动作流水线,利用标注引导的视觉提示——边界框标注同时标出可抓物体与摆放位置,提供结构化空间引导;动作层面不用传统逐步规划,而是采用 ACT(Action Chunking with Transformers)的动作分块策略,让机器人更稳健地完成抓放任务。 Robotic pick-and-place in convenience stores faces dense arrangements, occlusions, and varied object properties that complicate trajectory planning and grasping. This pipeline leverages annotation-guided visual prompting, where bounding box annotations identify both pickable objects and placement locations for structured spatial guidance. Instead of step-by-step planning, it employs Action Chunking with Transformers (ACT).
关键要点
- 便利店场景物体密集、遮挡多、属性差异大
- 边界框标注同时提示可抓物体与摆放位置
- 用 ACT 动作分块替代传统逐步规划
💡 对普通人的影响:暂无直接影响;无人零售与仓储自动化的落地成熟度有望提升。