Hi team,
Thanks for the great work on this project! I have a question regarding the current training data composition and pipeline.
I noticed that Visual Referring data—specifically, data where the query contains a bounding box and the response describes the content of that box—does not seem to be included in the multi-task training.
Could you share the reasoning behind this decision? Theoretically, incorporating this type of region-specific data should enhance the model's visual grounding capabilities and potentially improve overall performance.
I would love to hear your insights on whether including this data could be a viable direction for performance improvement. Thanks in advance for your time!
Hi team,
Thanks for the great work on this project! I have a question regarding the current training data composition and pipeline.
I noticed that Visual Referring data—specifically, data where the query contains a bounding box and the response describes the content of that box—does not seem to be included in the multi-task training.
Could you share the reasoning behind this decision? Theoretically, incorporating this type of region-specific data should enhance the model's visual grounding capabilities and potentially improve overall performance.
I would love to hear your insights on whether including this data could be a viable direction for performance improvement. Thanks in advance for your time!