- [2026.0705] 🔥 Our code is released!
- [2026.0522] 🔥 Our paper is up on arXiv.
We revisit visual planning in VLA systems and argue that effective planning should be local, visually grounded, internally generated, and directly aligned with action.
We propose Afford-VLA, a unified framework that internalizes task-conditioned affordance as an explicit visual planning interface, enabling interaction regions to be directly learned and leveraged for action generation.
See Installation instructions.
If you think this work is useful for your research, please use the following BibTeX entry.
@article{wang2026afford,
title={Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance},
author={Wang, Runze and Fu, Yuqian and Li, Yu and Lin, Tao and Qian, Tianwen and Elhoseiny, Mohamed and Zhao, Bo and Fu, Yanwei and Jiang, Yu-Gang and Xue, Xiangyang},
journal={arXiv preprint arXiv:2605.24203},
year={2026}
}
Thanks for awesome works: starVLA , RAGNet and Qwen-VL. Code is based on these works.


