Pith. sign in

Paper Citation Record · LEDGER

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2505.20728.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20728 v4

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:50:34.081389Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T06:34:56.032634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T21:11:15.085382Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c399578-31b0-470e-9f41-1b4bc570dda5 · outbound

This paper cites online" 'onlinestring :=.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.566175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.566175Z digest=sha256:f0a63567e42eef5d9312e2a57fc1f1533f05f5779abe20d56178da1b3f6b53ae

Observation 9d151fad-124c-448c-802b-279daa1da518 · outbound

This paper cites write newline.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.634571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.634571Z digest=sha256:17ef67c690a30fd3e8665523747d5f3792a1b26884622f5dfab1ff50506eb68f

Observation b7040f26-b49f-46db-9f98-0ba1569ab0bb · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.732806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.732806Z digest=sha256:f93144fb9e8ec81c3e2117acbea8ae520e1e34208d7bd12508228b3c6aac5bb3

Observation 752de2f1-8c9d-4fd0-b9d1-e72e290aab7a · outbound

This paper cites GPT-4 Technical Report.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.813020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.813020Z digest=sha256:48c9cfc86a1bf4475df26eda84408b34d758d069fff0f6c26bc713bcf8cf40c1

Observation 915d26d0-19a8-4d7f-8af0-09413637133e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.868504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.868504Z digest=sha256:d11ce5ee8a5a413df37424c6c1451e8c662a7c9900d258b1e87b333ceb718036

Observation a998c525-c809-401b-a93c-920b0ea44347 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.952646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.952646Z digest=sha256:c0e75af0ea02c9c276fbbfef7f8a37cb7c0cba57b3cfd5bdfb4685bf6573437e

Observation 5d9339a7-f16c-43db-b718-064548533b5c · outbound

This paper cites Qwen2.5-VL Technical Report.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.034463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.034463Z digest=sha256:bf5f21f810ab62d6da91d5b3bca09a022af9d13005eaf5c9552b581c4526b166

Observation 18f9b066-a05a-4a05-8efe-ba769070c060 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:36.299628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:31.115810Z digest=sha256:124cddbec21d1b813f752654feb5fff6b29f061a2e98e2fc821eba21a9737390

Observation 83229649-630f-4331-9b98-ad31b15475fc · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:36.156624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:31.165621Z digest=sha256:90fdcc70eafe85f744d805993727b6f2debce8f8eb25a98b782dc8f53a7f7fbb

Observation 0b6d4f18-972e-4d8c-b198-074d5dd7cb52 · outbound

This paper cites Aya Vision: Advancing the Frontier of Multilingual Multimodality.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Aya Vision: Advancing the Frontier of Multilingual Multimodality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.236941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.236941Z digest=sha256:02f28d046313d6ff7ec47fe2fafb8fe69d27ba7505e817e53b1b97bc224a46f2

Observation 88c75eb4-923e-4bb7-bbb5-8cb2c3f30022 · outbound

This paper cites Kimi-VL Technical Report.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Kimi-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.301410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.301410Z digest=sha256:b48f19b545218e8960fa60910b87cfac2e7e61ef69f02e266ab3b67c949c5f34

Observation d67f9f9f-cb3a-43ab-87b8-be3144fcc46f · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.970655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:31.377792Z digest=sha256:d5373e08e535fc04c0a47a299c79c65f870904c25b11d0ba4fad01a9c3bba29b

Observation 4b860dae-a052-4a8b-bf93-b134355876e7 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.444781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.444781Z digest=sha256:f5ed707861818382f778ae3263651ec12119c49f253043be7d5f6efe9e828325

Observation 1e2bdc44-4582-4ee6-866a-cbbed726afb1 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.779086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:31.516823Z digest=sha256:858e007bc812d22a7835a3d83d7123cd0617600af8f41d5d8a12897cb54dc9e0

Observation 56b1f494-4cca-4dd8-8989-6add8b57b1e8 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.590650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:31.595020Z digest=sha256:f940d57d43a183e4989f7ccc8d492abe6bd5aa4db0944cd12e5ad1f5582d3b3d

Observation 06300119-8468-4e4b-b66c-f7cbd0fe8a1f · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.686596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.686596Z digest=sha256:ac9ad90a4f68c16214838224316fdd6c74eed62a6514abd2aa41f890d54a71e4

Observation 8d8cc75a-6a49-4d13-b764-86b738695d19 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.769097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.769097Z digest=sha256:6e87dfb8c48d7cbf3d402399f155e61ab73136a6de630b390a3b11c157d50c38

Observation 1751a488-47f6-463a-95ee-3434fdd5a170 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.864485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.864485Z digest=sha256:9eb2b613fb380ba97ad7d092a5306eeb2613e9983954dc53042be24766d96087

Observation 98096c57-9bd8-444f-a72b-8ac62487f656 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.905490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.905490Z digest=sha256:46050db8261df6c04a312a105cd150b8f683fad2a0ed94fcab0fe182a8b946e4

Observation 8835969b-b7aa-46d7-ab34-527f0d1550ba · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.978776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.978776Z digest=sha256:16e167132bf207498cbdcd5818a73503a16f3a84baa6f4ad6f4ce08eb1c94a64

Observation e95c8114-bbb1-485f-a17f-6ce584870e4e · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.068720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.068720Z digest=sha256:e3d8e006a32be26e900ad40d2c38f461502907bb3fc1f9c023f89b94b4e81ec3

Observation 001d6cb3-83f1-49ba-98c9-d03602fda910 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.334929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:32.171204Z digest=sha256:09266d25828be18c1da499366feccc99191dc3231e9dd3043eb312a598d3468f

Observation 601b543a-3e8e-40a9-ba87-49474f2fca34 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.138981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:32.242418Z digest=sha256:f358c625a3b0391cc2d38fc98364db5982b488d065ec1d69f49f253f390d3161

Observation 0f972e68-e574-42ee-a6ed-21d647e6ff52 · outbound

This paper cites CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.315784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.315784Z digest=sha256:accf3c535f2b12c41405cc78f678a798def9bee21bc27fdeced25220e68c29ff

Observation ead65715-e649-4dca-b20c-e1f20b6c4f0f · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.032708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:32.389330Z digest=sha256:cdc9b8aad5261ed6acf9e52b0e70a752e693e2676a8f6f467db928c6e86a9f4a

Observation 5c97cc13-0618-4740-880f-af2c42d94a96 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.470962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.470962Z digest=sha256:275026cd9bfaccb5963bc8172c5b7b43f5d0d839a672fe5d5c1077b6c82bb7ad

Observation ffbe223e-d9c3-4ae9-9117-c5fa30c22be1 · outbound

This paper cites VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.535373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.535373Z digest=sha256:0d4fc5933a166999c6b8064bb09761af0320e89ecef50792b5d7d63a194fd6fe

Observation dfb006dd-2a19-4373-a98d-e6b997626f22 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.608633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.608633Z digest=sha256:3c5b4859147e56752769c3dada3e1f380db63e6be8ed8b98f2eaac8e4218171a

Observation 874f6b88-88ce-4ec7-a840-424a733f8385 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.676073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.676073Z digest=sha256:f5ee8e78aa20d4cc825c82216f4d3b6e8e18d4f8ab6660ce9f19a2e28ff414aa

Observation 4eae3f36-1e9a-4657-b9a4-c3b516b3d359 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:34.893955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:32.758845Z digest=sha256:83cc0e3a2b1df2d61059e1f91f2566c9c3937ffb4ecf78671aaf2360d785c62c

Observation 62794210-eb1e-45bc-b7fd-b271d7ac7dc9 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.851572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.851572Z digest=sha256:1c86fa52a1984d9c97efe83bd07f467e5a7e1db2807c324b2e6b46f9c7450993

Observation dc85e25c-5131-46e0-a594-a8fb0bacfe77 · outbound

This paper cites VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.927998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.927998Z digest=sha256:c480cdd9b9a84c7cac6f84dd4d41b88418b0b2a385a2fb528c2ef08574e325ca

Observation fc5d705f-b0e8-4de9-a506-d73312103c29 · outbound

This paper cites Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.010530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.010530Z digest=sha256:2645841dc24a46497a2ff37c8efa43cc5c6811720ec7c9d8ad59bc90928eb8b1

Observation f462880d-35a6-456c-a477-9290ac0be834 · outbound

This paper cites LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.120737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.120737Z digest=sha256:928dcd54e45734045da29bc1b774c5496c8cb09454f55e631032067e2a7ef0ee

Observation d5a87d20-629c-4fed-a2ca-1aa7cc8a503d · outbound

This paper cites Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.214128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.214128Z digest=sha256:89e486e8d79beca219c45581afb2941d60ef64a6482f4f04ab2535d3e2029ff2

Observation cf46d1fe-c7ad-4a60-ae0d-05f05bc2541a · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:34.743151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:33.357115Z digest=sha256:ae1f71f53af135d389fd472237b0f1e2de5b30db8c0a48cb705673475bef30ba

Observation 5594654b-632a-47ac-b866-0c18d131b687 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.477473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.477473Z digest=sha256:2347e651970893e9e95029f3c6d48b93e8e2ea7704bfb5aab81531d02f4f263f

Observation 537d74e7-7f34-4249-9156-1de218b34bdc · outbound

This paper cites CogLM: Tracking Cognitive Development of Large Language Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models CogLM: Tracking Cognitive Development of Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:50:34.334219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:50:33.577929Z digest=sha256:b0795521e70e39be907a6518119139d74233d879fc95b8c094b701463e05a2f7

Observation 1633c790-48b4-4a8f-ad8c-b011cab5840c · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.703965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.703965Z digest=sha256:2f7adedc2f4f003f18c8d3a75550d125deda4f0ec7b50a8c9f83ad3eef9b816a

Observation ae6ea074-f45a-4264-af88-2e2f6d623eb0 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.819280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.819280Z digest=sha256:f592a31800d2319ac3265c6ce0e86b4b7d4cee8a3b535b85dfbc3f7898b5e2ef

Observation 01d53ced-5ee7-4b5e-b343-a35ca024c4f8 · outbound

This paper cites Redundancy Principles for MLLMs Benchmarks.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Redundancy Principles for MLLMs Benchmarks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.938748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.938748Z digest=sha256:b5db51da4d72bb053dff8edce8b80a517c4074c21e0b5c8352fa937b67c9c9b5

Observation d9652d4f-5803-404f-bfbe-04cbc5d6017a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:34.081389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:34.081389Z digest=sha256:e0744bfd82d0d0d80e28161040d6b0ee7e518b659152cb9465139ed53c7cd53a

Pith citing papers

Observation af191167-4ede-47cd-87e2-f3dc88dfe650 · inbound

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction cites this paper.

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:15.133893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T06:34:56.032634Z digest=sha256:dc055b67d72e95f3c5e395edf5da43fd1032d92749d50c739b0872e20b02b196