Pith. sign in

Paper Citation Record · LEDGER

Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2310.00653.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.00653 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:24:58.819520Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 217ca226-394d-49d1-82e7-fb462604f334 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.309406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:0837220712b367df1f1f2f4f7877be51481a04246d3c0cbba2401d636c5d30c5

Observation 6aa60318-4518-4baf-b743-feb9d0bdf296 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.307160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:067d12674a7649403ea1517df8bcce80b8d649dae8bf7c877ec4a78fbc4d300b

Observation b5641bb9-40a0-4262-8502-9a954cfc4b04 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.087932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:075289c20338e820e504e2c37a502003af14b95a507041c43ed1244afc05d8aa

Observation 0b33969a-66ec-4006-8894-d29737ff13b1 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:42.993709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:d954a481376495f3bac933f4d6d4cfa00103297051d04c0bd54566cdd5fe35fa

Observation c745772d-88be-4c47-9f00-6d7516d1e0ac · inbound

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor cites this paper.

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:24:58.819520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:24:58.819520Z digest=sha256:71b4a50ef9ce9384fa67b8a7e0b2b1dac4cb34b22b751c41d5dbe37417842598

Observation 1edcaefa-fd54-4104-8ec9-14bc52363ac6 · inbound

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs cites this paper.

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T11:53:55.481346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:53:55.481346Z digest=sha256:3af3168dd07103e8ab6fb688ba91f2aaa7dbc4b43fc4948edd7b0b5a911f6936

Observation 82484be8-65d5-475b-8c38-226399ffc1f5 · inbound

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs cites this paper.

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T22:43:59.490358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:43:59.490358Z digest=sha256:791b68f3f7b819caf1308f26ceeefcf99ad1d9094af8e57a402e9cc86eff9842

Observation d24988dc-67bf-42b4-b1df-8f20cb1da434 · inbound

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning cites this paper.

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:59.580874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:22:59.580874Z digest=sha256:a0d93d77ddbdf05b626eba15b6ac5b1cb2108a7dd883d9e62ee700337434fd20

Observation f30a3137-a4df-42e0-871f-31751ce3591c · inbound

Deep Pre-Alignment for VLMs cites this paper.

Deep Pre-Alignment for VLMs Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:27:39.259198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T16:26:41.094936Z digest=sha256:7274ea7f5c8e121317c06ca43a57eef2d1dc30473e0c387a11a6a82636e14677

Observation 965a76ad-644f-4ffc-87e6-c643293715e6 · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:48.417157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:539ac294daee32b062f2ae04d939d188f2de15137ca7016a106c967e94eda405