Pith. sign in

Paper Citation Record · LEDGER

Visual Credit Audit for Multimodal Spatial Reasoning

As of 19 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.27069.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.27069 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:15:15.817164Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e0bfdb3-974d-4d23-8dd3-16b606184075 · outbound

This paper cites Answer the user’s question with exactly yes or no,.

Visual Credit Audit for Multimodal Spatial Reasoning Answer the user’s question with exactly yes or no,

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-01T10:15:15.817164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:15.817164Z digest=sha256:899f0db0b2522d39a913b7d8671888044717a1480e0540f6d095413320d15929

Observation 1168729e-5579-40e5-b34d-a050c68b2c76 · outbound

This paper cites VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning.

Visual Credit Audit for Multimodal Spatial Reasoning VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:13.651651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:13.651651Z digest=sha256:6a67f2b61e5dad2d67d20d5354767f57145f4f89ee7f53f0d6a74734ad01d21c

Observation 3b248b6f-ecdf-477f-89a8-990184a7a99e · outbound

This paper cites InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2901–2910.

Visual Credit Audit for Multimodal Spatial Reasoning InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2901–2910

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:13.884949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:13.884949Z digest=sha256:1c539c09914f44725c68e6b262ff050176f9127310f006b371c335c44e9944f7

Observation 08613a70-6534-43ac-805b-5d14fc85d25f · outbound

This paper cites From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA.

Visual Credit Audit for Multimodal Spatial Reasoning From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:13.957224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:13.957224Z digest=sha256:519d120da777c8ac57bdffe46fc772f5ca31776b4cf8b87130d98043682149f5

Observation 7f2c230e-7d58-4dee-82b8-89759e04746b · outbound

This paper cites Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?.

Visual Credit Audit for Multimodal Spatial Reasoning Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.125422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.125422Z digest=sha256:c17285ea2a5bdebf1260f483a068910f986eea2002fcbe8726e80bf29db27e7c

Observation eea224a7-5399-4fc2-b1aa-ddb2fa4920a9 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Visual Credit Audit for Multimodal Spatial Reasoning Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.442236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.442236Z digest=sha256:6de529fdd4aef7518f299689dd2e5722061923b90a83043c0a29ee10637f395d

Observation 98600053-1438-4679-9a6c-dc231412800f · outbound

This paper cites SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors.

Visual Credit Audit for Multimodal Spatial Reasoning SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.614006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.614006Z digest=sha256:6aa4b3dc745b9d23c08134b512f216d792c2b9fd03c48a02fdb0305b59c6a989

Observation cebd976f-5f52-466e-9950-90c4b60ec826 · outbound

This paper cites Ministral 3.

Visual Credit Audit for Multimodal Spatial Reasoning Ministral 3

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.765305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.765305Z digest=sha256:4eaad9f137febf686852924264b255b61c57b37b4904f3af580fb6955c2bcfdf

Observation d3f68a67-cd91-4ec1-9a5d-9f541d4ee56c · outbound

This paper cites GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs.

Visual Credit Audit for Multimodal Spatial Reasoning GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.867090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.867090Z digest=sha256:84d36095060c8714dec601f82dab5b693c9f1925ee6223969000cb0746f0be76

Observation 193d6c46-f062-4760-a173-b2c248592447 · outbound

This paper cites Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark.

Visual Credit Audit for Multimodal Spatial Reasoning Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:15.155395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:15.155395Z digest=sha256:a635dee453a2c4dcbaa16361dbd6999e581ee68a5c24350883af5350e8b32a8a

Observation 0aa24aef-0ac6-4617-a9cc-c80b05177381 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Visual Credit Audit for Multimodal Spatial Reasoning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:15.290623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:15.290623Z digest=sha256:110e9419b791c57bc563713c563ad3559b79131f32cff4cbdc6004c3d57bbbb5

Observation e1bb9bd4-c957-4048-83b3-0806ae941311 · outbound

This paper cites In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 10006–10030.

Visual Credit Audit for Multimodal Spatial Reasoning In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 10006–10030

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:15.617807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:15.617807Z digest=sha256:f5c8ad90835ae0849519977aa67ed748f39a1ff35206489cff1c8ddae6b99162

Observation f188f55f-56df-40ad-9229-32b1756dc62c · outbound

This paper cites Zhang, L.; Zhai, X.; Zhao, Z.; Zong, Y.; Wen, X.; and Zhao, B.

Visual Credit Audit for Multimodal Spatial Reasoning Zhang, L.; Zhai, X.; Zhao, Z.; Zong, Y.; Wen, X.; and Zhao, B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:15.755818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:15.755818Z digest=sha256:f511b2c4a6d259ab1f7b903c524d1205522380720211e3fb16488d45fcafed42

Observation c2086dd0-2351-4715-afb8-78ceb86bcb96 · outbound

This paper cites Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning.

Visual Credit Audit for Multimodal Spatial Reasoning Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning

Reference 305

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.321927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.321927Z digest=sha256:69af699377c73a85413a7b40d0f3d5d16ea44b642452091b7867f49d80837081

Observation 22d0b39a-7582-4bea-97f2-ddfd707ce5f3 · outbound

This paper cites In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6904–6913.

Visual Credit Audit for Multimodal Spatial Reasoning In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6904–6913

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:13.773475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:13.773475Z digest=sha256:8880e364d45d9b67466fa537abd8602b6d2e8da54d5bb1db5892545302b0bf38

Observation 72c2bec2-a363-4cf9-a891-f2409f179eb4 · outbound

This paper cites InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 4035–4045.

Visual Credit Audit for Multimodal Spatial Reasoning InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 4035–4045

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.996997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.996997Z digest=sha256:ebd79ecd11d7ac99c372636e494d97c1bda160c71d8771b0ea86ed42acda4dc4

Observation bb3fffac-79a5-45e3-b67a-4b6078f3cd32 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Visual Credit Audit for Multimodal Spatial Reasoning InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:15.472902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:15.472902Z digest=sha256:cf4495d1185d54881266b1e795651186e18c681a1b630ae31779b48b644725a4

Observation 218e5e27-d651-4439-8ef7-02f67f59a569 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Visual Credit Audit for Multimodal Spatial Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:13.210865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:13.210865Z digest=sha256:7c52944abd1c77882214faedec2b96265855d576526583100db67236e38a02a3

Observation 40f0d550-8e54-46f6-b874-2eac523f2d90 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Visual Credit Audit for Multimodal Spatial Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.197516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.197516Z digest=sha256:09f0488d0a731e383d896ead6a10d1087faf026f67ed0b8b16ac18092f716a92

Observation a4adac75-17eb-4a8e-b376-7290f0f57112 · outbound

This paper cites Qwen3-VL Technical Report.

Visual Credit Audit for Multimodal Spatial Reasoning Qwen3-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:13.302565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:13.302565Z digest=sha256:b41286944d68d8597c01cbf2aaa063779c48576988738e37095cbf26ad54dbe1

Observation 648f2f6d-4656-425a-8dcc-5e55166b7ba3 · outbound

This paper cites OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents.

Visual Credit Audit for Multimodal Spatial Reasoning OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:13.431093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:13.431093Z digest=sha256:cfa50d4d43f0277702911d23c9430a6716fca1af7bcab5c12d25363c9433296c

Pith citing papers

No inbound Pith citation observations are available.