{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:EFP2DXNGKEE7KAXDF2P2GW3VID","short_pith_number":"pith:EFP2DXNG","schema_version":"1.0","canonical_sha256":"215fa1dda65109f502e32e9fa35b7540d4ea2b42b8e66638c14695bbdcb02364","source":{"kind":"arxiv","id":"2503.23064","version":2},"attestation_state":"computed","paper":{"title":"VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models","license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Filippos Kokkinos, Junlin Han, Konstantinos Tertikas, Sabine S\\\"usstrunk, Shalini Maiti, Tong Zhang, Yufan Ren","submitted_at":"2025-03-29T12:50:38Z","abstract_excerpt":"Large Vision-Language Models (LVLMs) struggle with puzzles, which require precise perception, rule comprehension, and logical reasoning. Assessing and enhancing their performance in this domain is crucial, as it reflects their ability to engage in structured reasoning - an essential skill for real-world problem-solving. However, existing benchmarks primarily evaluate pre-trained models without additional training or fine-tuning, often lack a dedicated focus on reasoning, and fail to establish a systematic evaluation framework. To address these limitations, we introduce VGRP-Bench, a Visual Gri"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2503.23064","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","primary_cat":"cs.CV","submitted_at":"2025-03-29T12:50:38Z","cross_cats_sorted":[],"title_canon_sha256":"209330c47888af97b07edc4094805c91bf8059f19c76dd698efb5d945055ecee","abstract_canon_sha256":"f6b50b012468001ef21fcf28c0b5f4d3bb6faeee3212f735ae54d843a09e0e1a"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:43:07.082949Z","signature_b64":"yWCLnFdQdpAKi5C5H7gvayzAYT7+neOw9bTUZgS4Zi0XPhM8a8OLIeYh2wD21Z4f6frrJ/LMnGiHHLgFSNq8CQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"215fa1dda65109f502e32e9fa35b7540d4ea2b42b8e66638c14695bbdcb02364","last_reissued_at":"2026-07-05T10:43:07.082447Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:43:07.082447Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models","license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Filippos Kokkinos, Junlin Han, Konstantinos Tertikas, Sabine S\\\"usstrunk, Shalini Maiti, Tong Zhang, Yufan Ren","submitted_at":"2025-03-29T12:50:38Z","abstract_excerpt":"Large Vision-Language Models (LVLMs) struggle with puzzles, which require precise perception, rule comprehension, and logical reasoning. Assessing and enhancing their performance in this domain is crucial, as it reflects their ability to engage in structured reasoning - an essential skill for real-world problem-solving. However, existing benchmarks primarily evaluate pre-trained models without additional training or fine-tuning, often lack a dedicated focus on reasoning, and fail to establish a systematic evaluation framework. To address these limitations, we introduce VGRP-Bench, a Visual Gri"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2503.23064","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2503.23064/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2503.23064","created_at":"2026-07-05T10:43:07.082512+00:00"},{"alias_kind":"arxiv_version","alias_value":"2503.23064v2","created_at":"2026-07-05T10:43:07.082512+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2503.23064","created_at":"2026-07-05T10:43:07.082512+00:00"},{"alias_kind":"pith_short_12","alias_value":"EFP2DXNGKEE7","created_at":"2026-07-05T10:43:07.082512+00:00"},{"alias_kind":"pith_short_16","alias_value":"EFP2DXNGKEE7KAXD","created_at":"2026-07-05T10:43:07.082512+00:00"},{"alias_kind":"pith_short_8","alias_value":"EFP2DXNG","created_at":"2026-07-05T10:43:07.082512+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":10,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.19965","citing_title":"ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2606.19338","citing_title":"Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games","ref_index":65,"is_internal_anchor":false},{"citing_arxiv_id":"2606.08034","citing_title":"Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems","ref_index":65,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09883","citing_title":"The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2602.18600","citing_title":"MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2602.18600","citing_title":"MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11223","citing_title":"Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11223","citing_title":"Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09883","citing_title":"The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2604.16054","citing_title":"Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs","ref_index":3,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/EFP2DXNGKEE7KAXDF2P2GW3VID","json":"https://pith.science/pith/EFP2DXNGKEE7KAXDF2P2GW3VID.json","graph_json":"https://pith.science/api/pith-number/EFP2DXNGKEE7KAXDF2P2GW3VID/graph.json","events_json":"https://pith.science/api/pith-number/EFP2DXNGKEE7KAXDF2P2GW3VID/events.json","paper":"https://pith.science/paper/EFP2DXNG"},"agent_actions":{"view_html":"https://pith.science/pith/EFP2DXNGKEE7KAXDF2P2GW3VID","download_json":"https://pith.science/pith/EFP2DXNGKEE7KAXDF2P2GW3VID.json","view_paper":"https://pith.science/paper/EFP2DXNG","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2503.23064&json=true","fetch_graph":"https://pith.science/api/pith-number/EFP2DXNGKEE7KAXDF2P2GW3VID/graph.json","fetch_events":"https://pith.science/api/pith-number/EFP2DXNGKEE7KAXDF2P2GW3VID/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/EFP2DXNGKEE7KAXDF2P2GW3VID/action/timestamp_anchor","attest_storage":"https://pith.science/pith/EFP2DXNGKEE7KAXDF2P2GW3VID/action/storage_attestation","attest_author":"https://pith.science/pith/EFP2DXNGKEE7KAXDF2P2GW3VID/action/author_attestation","sign_citation":"https://pith.science/pith/EFP2DXNGKEE7KAXDF2P2GW3VID/action/citation_signature","submit_replication":"https://pith.science/pith/EFP2DXNGKEE7KAXDF2P2GW3VID/action/replication_record"}},"created_at":"2026-07-05T10:43:07.082512+00:00","updated_at":"2026-07-05T10:43:07.082512+00:00"}