Skip to content

Add ProVisE: Show, Don't Tell — Evaluating Spatial Cognition in Generative Pixels - #284

Open
zwq2018 wants to merge 1 commit into
BradyFU:mainfrom
zwq2018:add-provise
Open

Add ProVisE: Show, Don't Tell — Evaluating Spatial Cognition in Generative Pixels#284
zwq2018 wants to merge 1 commit into
BradyFU:mainfrom
zwq2018:add-provise

Conversation

@zwq2018

@zwq2018 zwq2018 commented Jul 24, 2026

Copy link
Copy Markdown

Hi, thanks for maintaining this excellent list!

This PR adds our recent work to the Evaluation section (top of the table, following the existing format):

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text (ProVisE)

ProVisE is a benchmark-agnostic framework that elicits protocol-constrained visual answers (pointing, marking, drawing) from image-generation models and parses them into structured predictions compatible with original metrics — enabling image-generation models and text-output VLMs to be evaluated under the same spatial task semantics. It also introduces SpatialGen-Bench, a diagnostic benchmark of 470 samples across 14 spatial subtasks and four capability levels.

🤖 Generated with Claude Code

https://claude.ai/code/session_018ecRPVPjWS3SK42jELrgBT

…-generation models

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ecRPVPjWS3SK42jELrgBT
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant