Pith. sign in

Paper Citation Record · LEDGER

Efficient Multimodal Learning from Data-centric Perspective

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2402.11530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11530 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:55:50.072148Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10fc5fcb-2c2c-449c-bf9d-7a35a5b434e0 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Efficient Multimodal Learning from Data-centric Perspective

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.158828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:78e455bfcb90552ed30c1cd05ef39cab919a1bd7f8f8ab7bc2c9043c75056af6

Observation 5dfdbce2-9602-4248-8841-61d285c8c824 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Efficient Multimodal Learning from Data-centric Perspective

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.058584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:1730159135631ce23d3084dbe930d1f51246e7f09f4da41deefecb25b2186ae9

Observation 452d07d0-b83c-43a4-a3bf-d1f59e28ffea · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Efficient Multimodal Learning from Data-centric Perspective

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.169541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:4794d238b7dff00289a0ec26e3549762ed0b6140f3a78d6a6b66833433bb2bac

Observation 2c8162b7-1d8f-42ab-954f-5ad194dd303f · inbound

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models cites this paper.

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:03.895366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:21:03.895366Z digest=sha256:dc51a249c226a5c83c30e261df9eab59ecf0d75f163b2a2e9dcf90642a1252ed

Observation 68ae5757-6f11-422c-a62d-ccb59f4426d7 · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:20.145924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:20.145924Z digest=sha256:8a3486a8bc29d0e62b0250057913c0218ebe18774d3704819e484963136d7db7

Observation 97b4e0f1-3f29-4140-b356-8d4085e4f5d3 · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.920866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.920866Z digest=sha256:5b07a8f988ae1005796b9b7ba36d473fc860091250c9233cdf88d1b07b9cea78

Observation e6e16ed1-ea5d-45d0-a58b-b8b29de3acce · inbound

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study cites this paper.

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:45.165347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:45.165347Z digest=sha256:37f441451d3538747fafffbe39f9009dd39beaacd0e19f2ae44ef131be3ff7d8

Observation b3ce6c97-49e4-41d6-96eb-d9643dce6b3b · inbound

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models cites this paper.

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models Efficient Multimodal Learning from Data-centric Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:35.876386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:49:35.876386Z digest=sha256:324d23d300b3d7be57267ac970977cc4f9d1514f75d76601bbfefba377586ddc

Observation f088ccac-837a-4e8c-bd5f-478b8ee0fe35 · inbound

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings cites this paper.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Efficient Multimodal Learning from Data-centric Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.894715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.894715Z digest=sha256:66acaf4eadd5c85d1cb31ab9ca87025fc0d9041968fc7208867e7d7b47125147

Observation 82da0299-3d75-4c2d-ab49-b347be8f59b6 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.053118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.053118Z digest=sha256:6bed912b8c776ca92a5e9fc30753467927cb1227ee637f0f2e370eebefbc85ef

Observation 453a9f6f-2842-496c-8795-b92fff2e9f58 · inbound

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models cites this paper.

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.117667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:17:29.924860Z digest=sha256:c865b147513bf751db5c0979b1bbec3e8e5da4a329d5d1d2fe09b47b95d1471d

Observation 53763547-4478-4476-bde3-62e380ba8eca · inbound

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models cites this paper.

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:34:27.393703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:34:27.393703Z digest=sha256:447ec58ee305e600ebc9fc72ef96a5a2484e2962f8eb13b01e7d76db2c83a897

Observation 5f13addc-ae1c-4578-a725-6069109847f4 · inbound

Leaderless Collective Motion in Affine Formation Control over the Complex Plane cites this paper.

Leaderless Collective Motion in Affine Formation Control over the Complex Plane Efficient Multimodal Learning from Data-centric Perspective

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T09:18:06.363249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:18:06.363249Z digest=sha256:3b635efd9aeb75e805d715e4fa936ae47d5db3659048bdbb2bcf7eedef250a1c

Observation 079bb065-94a4-430c-97a7-a60729d7dc27 · inbound

Analogical Reasoning as a Doctor: A Foundation Model for Gastrointestinal Endoscopy Diagnosis cites this paper.

Analogical Reasoning as a Doctor: A Foundation Model for Gastrointestinal Endoscopy Diagnosis Efficient Multimodal Learning from Data-centric Perspective

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:15:51.452284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:01:47.592716Z digest=sha256:dbbbca94d6353507a6f1e4875418adae50a298e94ee5a8664094b208441ad540

Observation e7156161-bfb2-415a-a446-535bd84dbd21 · inbound

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency cites this paper.

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:56:14.224287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:47:38.100037Z digest=sha256:ab181497abda71540d897c94dd5c27abaddc95ff2d82a52d15787311f3da36a5

Observation bb565ed3-7b21-4339-a475-b697be2ce694 · inbound

Anisotropic Modality Align cites this paper.

Anisotropic Modality Align Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.270268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:08:04.686087Z digest=sha256:ffde67283fe97e9ba4c7cfbc3bc85272f2d10004f0b5532a002a38e59bf81784

Observation 94ba5713-8354-48f0-8240-1361bc146040 · inbound

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models cites this paper.

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:25.660081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T05:10:56.991368Z digest=sha256:94dc4e773ceb4d65866c426afc34b888aade2c7251b530c1cc5d5942de146d69

Observation 7fa6a991-7251-46d8-b2d6-952b5690e897 · inbound

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens cites this paper.

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens Efficient Multimodal Learning from Data-centric Perspective

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.521871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:07:12.965100Z digest=sha256:3d8caf9f22b0be6176ae96e02b8022e2a5ff051411aeb60208d60f8cf9e77858

Observation ca881d48-ae33-4ddb-9937-4b00bb06aebe · inbound

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens cites this paper.

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:09:38.228602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:07:12.965100Z digest=sha256:8d2df5a370cea3b0967e6abfbab306c655c6767ec60d0163cc7203bb35cd6ff7

Observation 252bcedb-73af-4ce7-9495-d38b11e74add · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision Efficient Multimodal Learning from Data-centric Perspective

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.009876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:0cd1e40c5c276341a7abc162073059e675e9d6a23efda074175161de42a8c031

Observation a4f72ba6-9734-4009-a020-ac34093c6fbe · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning Efficient Multimodal Learning from Data-centric Perspective

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:48.414077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:f4a64995fcbed764ce7c9fad9a21704051b4ee15963935440c8110199f10f0e4

Observation 7d997c3f-dce5-4477-9ff6-d92ba060c17d · inbound

Bridging the Modality Gap in Forensic Image Retrieval cites this paper.

Bridging the Modality Gap in Forensic Image Retrieval Efficient Multimodal Learning from Data-centric Perspective

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.986808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:06:22.362644Z digest=sha256:c3dad2cbc93b789f660edc8082a455065a2304891704ea6cfbba268a9616d797

Observation 7f43068c-9afa-42d8-9621-b03a69f8ea93 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Efficient Multimodal Learning from Data-centric Perspective

Reference 278

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.477800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:bd6b327ac493c4c517ad4a2eedeaaeca294819883db088d3cf59849ac4b882ad

Observation 6edb94b7-7acc-4fc2-b394-41b60f279ce5 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Efficient Multimodal Learning from Data-centric Perspective

Reference 278

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.576159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.576159Z digest=sha256:e155930557d2e9d9df3a0d2bbed02d752659f49387263a482b08d5a04d835504

Observation 1373dba2-dee2-4b26-8797-591687a27972 · inbound

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing cites this paper.

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing Efficient Multimodal Learning from Data-centric Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T16:07:55.017604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:07:55.017604Z digest=sha256:3656d51d4c577d121fb9cbeeadf495fc919a237ec675e33a667dc85b8d061476

Observation 7e115841-d1f7-4f52-aae3-970b7392d232 · inbound

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models cites this paper.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.072148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.072148Z digest=sha256:702c65d828bf1ed84af6578afab164d6ce33de35bf6ab4d4445fa3a61cc18334