Pith. sign in

Paper Citation Record · LEDGER

Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2403.19322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.19322 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:28:10.614055Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.150854Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 521ffd76-dec6-4ecb-b8fd-ed6910b5d136 · inbound

Task-Core Memory Management and Consolidation for Long-term Continual Learning cites this paper.

Task-Core Memory Management and Consolidation for Long-term Continual Learning Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:28:10.614055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:28:10.614055Z digest=sha256:a250b3c4add223e9b4c24f29f939a07ad917f491dabefc23363e70db5680190a

Observation 987437d3-de0a-4659-b914-f0228fce6e58 · inbound

Explain Before You Answer: A Survey on Compositional Visual Reasoning cites this paper.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.825678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.825678Z digest=sha256:ed6eb79ffeae0fa4c7343c11e188a4baec8233d8a69db418a97fc71b44043198

Observation c7cce8f4-d3ce-4f3f-9164-84a4c427e005 · inbound

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning cites this paper.

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.687463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T22:38:30.213948Z digest=sha256:d0e4dc2e4e640ba12f6d623bac92ea2b5848803f2ca484abeacc80fb320be27f

Observation cf471cc0-cc72-4db1-b417-3885a11bfe6a · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:22:37.725378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:f4c5219df481ceb9402365f639216fb98bb6e26cd191f198cd6b9d89a93260ce

Observation e25c475f-01fd-4e84-ae2a-af11ab0be706 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.153014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:bc6ab8065e0a3c46d4ef4ae4ee4ed054882e42c232ae66114b5363836bf42028