Thank you for your excellent work! I would like to know if you have provided the Referring Image Captioning (RIC) dataset? If the processed dataset is used, it is only suitable for the QWEN 2.5VL architecture and not for other architectures. This will be very helpful for us to widely use your model. Thank you!
Thank you for your excellent work! I would like to know if you have provided the Referring Image Captioning (RIC) dataset? If the processed dataset is used, it is only suitable for the QWEN 2.5VL architecture and not for other architectures. This will be very helpful for us to widely use your model. Thank you!