研究突破 arXiv cs.AI
对话式图像编辑的下一步操作推荐研究 What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
精读摘要
对话式助手越来越多地推荐后续编辑来帮用户延续任务,但现有系统主要面向纯文本交互,图像创作场景被忽视。研究团队从 Qwen App 收集了 10 万条真实多轮图像创作对话,发现 80.1% 依赖图像内容,说明多模态推荐是刚需。图像场景的后续编辑建议需要同时满足三个条件:符合用户偏好、方向多样、且能在当前图像上真正执行。 Conversational assistants increasingly recommend follow-up edits, but existing systems target text-only interactions, leaving image creation underexplored. The authors collected 100,000 real multi-turn image-creation conversations from the Qwen App and found 80.1 percent are image-dependent. Useful image edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image.
关键要点
- 从 Qwen App 收集 10 万条真实多轮图像创作对话
- 80.1% 的对话依赖图像内容,多模态推荐是刚需
- 好的建议需符合偏好、方向多样且可执行
💡 对普通人的影响:使用 AI 图像创作工具的用户,未来可能获得更贴心的「下一步」编辑建议。