Pith. sign in

Paper Citation Record · LEDGER

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2604.20665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.20665 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T10:35:50.838150Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:51:18.755546Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact6
  • verified fuzzy4
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 68af16fe-7300-48c9-9cd7-9ce6eaa3f56f · outbound

This paper cites 2025.FastVLM: Efficient Vision Encoding for Vision-Language Models.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm 2025.FastVLM: Efficient Vision Encoding for Vision-Language Models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:36:25.070848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:020d7e02ba79dc4bf765e4fb0733dfb445aebc97079d8d06495ca377216cd9da

Observation d795f7b7-fd97-412b-8678-9d9a03b447c6 · outbound

This paper cites An Introduction to Vision-Language Modeling.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm An Introduction to Vision-Language Modeling

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:36:24.996880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:29302d52b9f764b220111941eb72caede9c5cdee3d6cc643dd2608f45bb61372

Observation 445a12ad-6f67-4c50-adec-386e139a9b8a · outbound

This paper cites an unresolved cited work.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-22T10:36:25.081898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:0d8bbd44ce361eec0493ade595742cbd1934b8caebcefe34359e8f8d34601f28

Observation 0eb02fec-f321-4728-bf03-db7c1a25d8e8 · outbound

This paper cites BabyVision: Visual Reasoning Beyond Language.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm BabyVision: Visual Reasoning Beyond Language

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:9fca8cd77a2f41e947ff79dc370c13d26f22dcbe11f1bd6174d0f1ad93be2ac5

Observation 4591128f-01c6-41d4-91b5-7adcbe14be72 · outbound

This paper cites Improving Fine-grained Visual Understanding in VLMs through Text-Only Training.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:36:25.017520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:2444d5ec6b2b1d1a894b164a172562eb97fa3a91595cfa99f9f52d6d60371ee8

Observation 146b8b59-4bc1-455b-8af5-e3b85825f028 · outbound

This paper cites an unresolved cited work.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-22T10:36:25.077967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:35c56ddabd0262201296a6fe8922c7e5a9f9ec1947d8db78b1a9b91db46dbd9e

Observation 61fbecdc-c32f-4da7-ad7f-aee2dcc69bdd · outbound

This paper cites an unresolved cited work.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-22T10:36:25.074389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:f0bc0d71174ca1920cb872b5f370f29f756c6e29cbe09e9986dae7d73764a49c

Observation 023d4941-62dd-4f22-8cef-16386b677f32 · outbound

This paper cites an unresolved cited work.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-22T10:36:25.067160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:3ec79cdfc85910d6b98e0a0b10150feecdffd1d14fa540264002290d7d9fbdd1

Observation 313e0342-0185-4ab2-912d-4c8579ffe582 · outbound

This paper cites 2024.What are vision language models?https://www.ibm.com/think/topics/ vision-language-models.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm 2024.What are vision language models?https://www.ibm.com/think/topics/ vision-language-models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:36:25.063602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:29bcfaa3a09937b9ac1ed39a1ed7cc97dc175f87377ba2e293c2a81a7e848478

Observation b1fb141f-8865-4542-a2d2-af25d41ca59a · outbound

This paper cites Explain Before You Answer: A Survey on Compositional Visual Reasoning.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:ea0046ca18e023c37207af05f72ef576f97ac63b2454088450107b9fc75baa76

Observation 088f0cb2-9736-469c-86f7-3c32d940cdfa · outbound

This paper cites an unresolved cited work.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-22T10:36:25.055505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:fcabcc67fb79584d9a9a375ee32d3805ea10f63d3ddd534ec26e70497b5a3af7

Observation 14baa9d3-77c1-4588-b95b-54671a85b9da · outbound

This paper cites 2025.Vision-Language Models.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm 2025.Vision-Language Models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:36:25.059410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:04f3012fa35f49f95a963d6880a84cfee00f6a159ba158658227422f83e11fa7

Observation 4c2cf8df-b5c4-4666-b70c-b238707de06c · outbound

This paper cites 2025.What Is a World Model?https://www.nvidia.com/en-in/glossary/ world-models/.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm 2025.What Is a World Model?https://www.nvidia.com/en-in/glossary/ world-models/

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T10:36:25.051573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:23457d914162d3073589e18f91ea3221b6434c7753f2ee7fa3375397fef890e3

Observation 5bf38efe-c0b3-4175-ae13-d0f1cb80751c · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:36:24.988618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:85f1e88e1fa409035c70717fafa325a84ece5ddc9d8bd01971713f63c6810460

Observation 88f4b28a-c779-4e96-a29e-3486d318ec72 · outbound

This paper cites an unresolved cited work.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-22T10:36:25.043190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:80286e0366edcd559c8ef1753161c2bf226c618ee6d737bb99288c4f8e96f10d

Observation 0e6f627a-d623-4693-88cf-97ea49aaafcf · outbound

This paper cites Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part VIII , pages =.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part VIII , pages =

Reference 16

Resolution
verified exact
doi, observed 2026-05-22T10:36:24.936103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:c5a487b826eba3ca84faa8fd31f6daa1ccff1a83b6c7b11104070113c9af7ae9

Observation 00354cf3-3d68-4623-a3ca-be829883cfec · outbound

This paper cites an unresolved cited work.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-22T10:36:25.047120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:7722f9494490cb0eaead7c4c0b1d4596db507dac0940e24178e17552ba54d57a

Pith citing papers

Observation ef11b8fe-3dc7-48cc-8887-094819e70512 · inbound

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models cites this paper.

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T10:51:18.755546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:51:18.755546Z digest=sha256:fa4238d577fd681c7393bd3bedda34da1e50311f33106cefe505dd28b725241f