Pith. sign in

Paper Citation Record · LEDGER

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2505.05464.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05464 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:15:06.746182Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df9585c7-f145-4952-9ca8-d51d77c80e69 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.876272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:ca78b53f8b1b7fd4000b90d244fec7db4962a4d265f08b1640bda08c70ed1f2a

Observation 52166a51-173b-4065-baed-1b8849038595 · inbound

PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models cites this paper.

PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:50.323121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:50.323121Z digest=sha256:44d021bfd6da054510c39c59ed1736fae84fbb4ec218facbc8add39ae4bf9a0f

Observation c9568ce4-e1e1-418e-97cd-37470082b171 · inbound

Training-Free Reasoning and Reflection in MLLMs cites this paper.

Training-Free Reasoning and Reflection in MLLMs Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:55.819405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:55.819405Z digest=sha256:4abdb4d163c48c46815b2ccd5f739bd4b2c9356cee24489408e9cf18c60bd2d3

Observation 18bcdaae-1d8b-42aa-9472-b489a4c8ac21 · inbound

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models cites this paper.

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:35.625880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:35.625880Z digest=sha256:ad94c01c433a7c62dc6fbbcc8d2e7bde04df29583365b91c9bcdda3a2dcdd931

Observation 3c15fbc6-41d7-49d8-9fe3-6b4c95398690 · inbound

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging cites this paper.

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:05.779148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:42:43.948462Z digest=sha256:f50ae125a0d069687b24490080b732a755aa493f7ca005c42cc2a26f9a6a5758

Observation 5e991e75-8394-4b31-8ba6-d69525255154 · inbound

Towards Long-horizon Agentic Multimodal Search cites this paper.

Towards Long-horizon Agentic Multimodal Search Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:00.190608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:40:32.137708Z digest=sha256:7a0fb136cbc448ca6df2c6a47cea87ef01bf88a22a0557a16f158f0584d3fc79

Observation dd4b5312-f3f0-48b6-aeb6-f3645440beb6 · inbound

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails cites this paper.

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.487917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T23:03:01.151745Z digest=sha256:4ba772b5e498f3f5e691b3f3c195d0a38d587d3399226d803cca7f9695368c35

Observation f9a56e1a-b4b3-4e7f-852a-9ff13eb396d9 · inbound

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning cites this paper.

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:06:16.623754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T15:45:26.891621Z digest=sha256:05e853247f390fe91a9d34cb441c34f3e1c93a7a97e90f5f0380fdbf60d43318

Observation 613f2f3a-19a1-4ec1-a37f-867336c1a40e · inbound

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models cites this paper.

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:56.348287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:10:12.974244Z digest=sha256:a57e7c0a85a949b327586e148cbcd41db40250138e6266aaca7ef91923573072

Observation e4417778-bbfb-4451-8ae9-3e4b983a85e3 · inbound

Recursive Vision Language Models for General Symbolic Reasoning cites this paper.

Recursive Vision Language Models for General Symbolic Reasoning Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:11:24.638371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:11:24.638371Z digest=sha256:2d1e76370f9ea0a262e508343297cb1ed88cf8462043a637337a458adb5556a6

Observation 97c0e90a-481d-4511-b135-8327fc93146a · inbound

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging cites this paper.

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:06.746182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:15:06.746182Z digest=sha256:781667a7888b2051862dd51ee8c59cd0b76db91f321064b21946a7a02e7ee552