Pith. sign in

Paper Citation Record · LEDGER

Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2310.00653.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.00653 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:24:58.819520Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 217ca226-394d-49d1-82e7-fb462604f334 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.309406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:af8307929500d420c161e562f9921530ca38ae988b80eec1160f37fa72825509

Observation 6aa60318-4518-4baf-b743-feb9d0bdf296 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.307160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:530ccf2af6b10db1708e675f920da80d4eff6f40b9dcc69910398a055a691445

Observation b5641bb9-40a0-4262-8502-9a954cfc4b04 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.087932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:9c327064dcc140f29f04823af924f218532bd1d6e3125c4f6a33f133b236b34c

Observation 0b33969a-66ec-4006-8894-d29737ff13b1 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:42.993709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:9c6d2214728bd2e0ce4600c76ed7e2de5d1c2f07c4ba86d2850f9785afef0509

Observation c745772d-88be-4c47-9f00-6d7516d1e0ac · inbound

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor cites this paper.

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:24:58.819520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:24:58.819520Z digest=sha256:a596c113b2a6332c6ed8f5f573e5933a8e9e66bdd8e47b9241e790dbf92aeb63

Observation 1edcaefa-fd54-4104-8ec9-14bc52363ac6 · inbound

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs cites this paper.

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T11:53:55.481346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:53:55.481346Z digest=sha256:1bd0d3413b4982fe1a05a884d2eb4ad4306b3402508ad078b381e09c691e0353

Observation 82484be8-65d5-475b-8c38-226399ffc1f5 · inbound

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs cites this paper.

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T22:43:59.490358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:43:59.490358Z digest=sha256:880b7c9fb831b49209e23db956d4f20e4aa4866ef4ef323176d55d41372f9dc4

Observation d24988dc-67bf-42b4-b1df-8f20cb1da434 · inbound

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning cites this paper.

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:59.580874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:22:59.580874Z digest=sha256:213ec9b3a5ccc29bbb915f3c94a24004a0c5a7e7956d08c0645e7e42373c3b7c

Observation f30a3137-a4df-42e0-871f-31751ce3591c · inbound

Deep Pre-Alignment for VLMs cites this paper.

Deep Pre-Alignment for VLMs Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:27:39.259198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-19T16:26:41.094936Z digest=sha256:86c720fe0e472ab5c65043e1aca490bba87f1260f59c3ae518e1dec6e033224b

Observation 965a76ad-644f-4ffc-87e6-c643293715e6 · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:48.417157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:954e324ac308372e6c3ecde26c28452f34c60c4ebd4c1ebecafa05fcc98a7131