Pith. sign in

Paper Citation Record · LEDGER

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap

As of 23 July 2026, this Paper Citation Record lists 11 of 11 outbound references and 2 inbound Pith citation observations for arXiv:2604.16256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16256 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T09:01:35.425918Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T07:43:50.633779Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-13T07:47:33.319161Z

Reference resolution

11 of 11 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9d5d633-9a7c-456d-9a8e-76e60b701bc8 · outbound

This paper cites OpenAI GPT-5 System Card.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap OpenAI GPT-5 System Card

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T09:03:25.041433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:1963ae5e9656705fe7e411768ca7f91a501c9e12816d23ce47b514739575cd81

Observation 4f39df01-7084-46e2-8517-4267c377a059 · outbound

This paper cites EasyARC: Evaluating Vision Language Models on True Visual Reasoning.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap EasyARC: Evaluating Vision Language Models on True Visual Reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:25.045165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:db13643f66a1970a83357585056592b491ccf4cd59b6294759be5698ce0ec9f9

Observation 86b2b7ce-77b7-49d2-80a1-a10ee83e03a1 · outbound

This paper cites Emogen: Emotional image content generation with text-to-image diffusion models.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap Emogen: Emotional image content generation with text-to-image diffusion models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:03:24.517640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:66946cdecea5ddf86e3dfc7c332d3ff3f03ad117ea08898a36f6f0e62a161aeb

Observation c8299b56-70f6-4268-af4e-2416ece439b4 · outbound

This paper cites URLhttps://doi.org/10.1007/978-3-031-73242-3 10.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap URLhttps://doi.org/10.1007/978-3-031-73242-3 10

Reference 4

Resolution
verified exact
doi, observed 2026-05-10T09:03:24.510527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:9fe57fa3108c6f1c356113f42977a9e35122987c4ae3e60ddb6b9c2f3b8b9f0d

Observation 03aa5c2d-3571-4e31-b762-40cedb41e5df · outbound

This paper cites You are given a **cross math puzzle** in a **textual markdown grid format**.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap You are given a **cross math puzzle** in a **textual markdown grid format**

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T18:23:38.173700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:89588be3472235116deaa1e04f5c54e89b67be7811e42971d9b12e97d56e1fd4

Observation c43f78a9-55c5-4b25-84e4-7c2adf8cc0d8 · outbound

This paper cites You are given a **cross math puzzle** in a **image format**.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap You are given a **cross math puzzle** in a **image format**

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T18:23:38.161128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:835fd7e15259dad52d42da9a97737d9360501db103089493440409567af1d99e

Observation b6a6c397-ee48-4675-9520-29d59af3fc58 · outbound

This paper cites You are given a **cross math puzzle** in **both image format and textual markdown format**.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap You are given a **cross math puzzle** in **both image format and textual markdown format**

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T18:23:38.169302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:d0e9d2a7232aecdc16e68aa9b85987e6921335f7f56b99f79ee519301c3683fd

Observation 56fad0d0-b3f6-4e86-bb69-4f921626dc61 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap Unresolved cited work

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T18:23:38.175408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:c187707aeeb4a4af27f304e89fd88ec2d5072804cd1d8bbc78fd2d0920426a9c

Observation 86efefc3-01d1-40c7-ac84-e50071f55c7d · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap Unresolved cited work

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T18:23:38.171322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:bd3e2282e4b92027a53e3d46db04582413909609aa5df2cf8f6064b6b083937b

Observation 1b155ff9-751c-4abf-b921-151e36394255 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap Unresolved cited work

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T18:23:38.163408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:f939a94378a52bda2f7c10d3cbfd91fee0bdc6a1743a841f0a0081a15df10ebe

Observation ac2646c3-2185-4d7c-8750-61950c38d326 · outbound

This paper cites You are given a **cross math puzzle** in a **textual markdown grid format**.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap You are given a **cross math puzzle** in a **textual markdown grid format**

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T18:23:38.166227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T09:01:35.425918Z digest=sha256:67047c53f2bed1afca6e9841aa260e697b2b6661f0b42196a18ce803ae2988e5

Pith citing papers

Observation bb1354e6-797d-4137-88e5-ff5f2f9e9713 · inbound

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning cites this paper.

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:36:26.064995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-12T05:07:39.571227Z digest=sha256:5957e98e11ca5caff789fc8f93a7c3b78f7ea05d7e38e710f77e18dc24df394e

Observation a8d8672a-8a01-4291-b1b9-c2915943612b · inbound

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning cites this paper.

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:47:33.322107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-13T07:43:50.633779Z digest=sha256:ccfeac526c78db681004dae05aaa3a019be804341eaf332401b3dc4642088e2f