Skip to content

Repository files navigation

Afford-VLA logo

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance

Runze Wang*,    Yuqian Fu*,    Yu Li,    Tao Lin,    Tianwen Qian

Mohamed Elhoseiny,    Bo Zhao,    Yanwei Fu,    Yu-Gang Jiang,    Xiangyang Xue

arXiv

🚀 News

  • [2026.0705] 🔥 Our code is released!
  • [2026.0522] 🔥 Our paper is up on arXiv.

💡 Features

✨ Effective Visual Planning Paradigm for VLA Systems

We revisit visual planning in VLA systems and argue that effective planning should be local, visually grounded, internally generated, and directly aligned with action.

Afford-VLA teaser

✨ Afford-VLA Framework

We propose Afford-VLA, a unified framework that internalizes task-conditioned affordance as an explicit visual planning interface, enabling interaction regions to be directly learned and leveraged for action generation.

Afford-VLA framework

📦 Installation

See Installation instructions.

📂 Data

See Prepare Training Data.

🦾 Training

See Train Afford-VLA.

📊 Evaluation

See Evaluate Afford-VLA.

✍️ Citation

If you think this work is useful for your research, please use the following BibTeX entry.

@article{wang2026afford,
  title={Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance},
  author={Wang, Runze and Fu, Yuqian and Li, Yu and Lin, Tao and Qian, Tianwen and Elhoseiny, Mohamed and Zhao, Bo and Fu, Yanwei and Jiang, Yu-Gang and Xue, Xiangyang},
  journal={arXiv preprint arXiv:2605.24203},
  year={2026}
}

🙏 Acknowledgement

Thanks for awesome works: starVLA , RAGNet and Qwen-VL. Code is based on these works.

About

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance

Topics

Resources

Stars

20 stars

Watchers

1 watching

Forks

Packages

Contributors

Languages