Pith. sign in

Paper Citation Record · LEDGER

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 3 inbound Pith citation observations for arXiv:2412.12940.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12940 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T21:20:00.041277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.415452Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f4f61ac8-4882-48ed-8eae-3bf6c8df07b5 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:51:05.108698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T23:56:02.856878Z digest=sha256:957a335278ed1a15bb7ee34bab0ebee653277a08001ea8eafef722d5d3815fe8

Observation 4591128f-01c6-41d4-91b5-7adcbe14be72 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:36:25.017520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:52921df36541e40002682a8df01613407916f249b8cbe61dd76f397e11912dca

Observation 8786e51c-aef8-47c0-9dcb-f02a802a4832 · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.416837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:6cce4e4c6ddeaea27d3ec732590163a4ca91203a2a29b510fbf2fd4bb8f12483