Pith. sign in

Paper Citation Record · LEDGER

A Review of Multi-Modal Large Language and Vision Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2404.01322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.01322 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:36:15.479733Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 71f318e2-1579-4182-b768-4d14b299a69d · inbound

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models cites this paper.

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models A Review of Multi-Modal Large Language and Vision Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:15.479733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:15.479733Z digest=sha256:0457eaba688432be5db4b978f32ea6281143273937355755e4bb311f83d6ff2f

Observation 9830cf20-e016-45ed-bdd3-d6c84b3cfadd · inbound

HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs cites this paper.

HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs A Review of Multi-Modal Large Language and Vision Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:41.097111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:41.097111Z digest=sha256:eee505166cb69d8a8a4707d0a6a9f4f3672c5364e439b6a35a0239a603eb1fa1

Observation 473f1fd4-803c-41fd-8dc1-524eea4a2891 · inbound

LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents cites this paper.

LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents A Review of Multi-Modal Large Language and Vision Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:48:43.288927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:48:43.288927Z digest=sha256:17a4e199ad655e37c35d998c58061349faa821a2f82141815309d985931a0317

Observation a4878e1b-08ea-4c31-b93e-a8c591c66d13 · inbound

Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering cites this paper.

Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering A Review of Multi-Modal Large Language and Vision Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:35.364167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:35.364167Z digest=sha256:8f672f78ca57ef1628457972090fe0d50b973f64e803ddfed8cabee9d705d7d2

Observation bf1612b3-304e-483a-84bf-ea8147f3fb1d · inbound

WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring cites this paper.

WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring A Review of Multi-Modal Large Language and Vision Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:10:09.177122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:08:34.674972Z digest=sha256:edde3172ff6e49cd0ea27d7581e7be4694a8ec1d3673d241d6b6d5d1f1b667b2

Observation b10e73ee-6c9e-4744-9d6e-7b6a957cd787 · inbound

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips cites this paper.

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips A Review of Multi-Modal Large Language and Vision Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.685148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T15:58:52.425995Z digest=sha256:f8046bbaf41009544ed636579e01b0d16e9290d709000901d60458802150f706

Observation ef9fd723-9775-4162-b6dd-aa42fe10423c · inbound

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips cites this paper.

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips A Review of Multi-Modal Large Language and Vision Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T21:58:47.248139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T15:58:52.425995Z digest=sha256:c910ffe6d25e6b1d8e57b0f9bbc1542bb096c2a686f33a7cdcdede44e22fbeab

Observation 57d959cc-61c0-4e1c-9a52-f17246a94efe · inbound

Rethinking Video-Language Model from the Language Input Perspective cites this paper.

Rethinking Video-Language Model from the Language Input Perspective A Review of Multi-Modal Large Language and Vision Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.475273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T13:24:46.360149Z digest=sha256:0e5bcfffe6a5cfb539d9bf049642c7639c7d96aca1595ed8733ea10621d40456