Pith. sign in

Paper Citation Record · LEDGER

Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2405.15973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.15973 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:39:18.941061Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fda87008-9453-4592-8f8a-c60a3b90b76f · inbound

Continual SFT Matches Multimodal RLHF with Negative Supervision cites this paper.

Continual SFT Matches Multimodal RLHF with Negative Supervision Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.111187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.111187Z digest=sha256:2570b783b64aa42bcd7f79dcca1250c44677044de36c2781bc7b065bc4475806

Observation 17cfad8e-0140-4bac-9f76-730c358064b0 · inbound

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning cites this paper.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.342823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.342823Z digest=sha256:93a4e40ef549c13ff4e25505ae808d39daa498174ed30265fb23495db1166498

Observation d30d8293-00ff-4c4c-b6e8-13062621e2a1 · inbound

VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning cites this paper.

VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:50:01.390763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:50:01.390763Z digest=sha256:e23cc3aebd444f2e19873816bad6326fb55c4da45aef2101940054f05bd9881d

Observation 9771fe61-2df7-4e0a-b422-aa1e4f6af4bd · inbound

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension cites this paper.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.663857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.663857Z digest=sha256:d54d947396ba1039e57159c1bdb79c45bad4d974b9f13aab5ee753ea4ea4858b

Observation b781e7cd-e719-4fb8-bdf5-c8ba61f9eb03 · inbound

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation cites this paper.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.422147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.422147Z digest=sha256:bb2536bec88e8db4a02724af92d0639123e7df23f142eb9dcb3bd37e17200ae8

Observation 12356855-ccbb-4ccd-a098-0106a4368fcd · inbound

Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution cites this paper.

Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:36.312215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:36.312215Z digest=sha256:4e233f70f0fda9a862d7d3d31358edc938b276b2d07da130433197c8e8c48010

Observation 25f84572-21ab-4133-ad66-cef7406e7eb8 · inbound

Probing Visual Language Priors in VLMs cites this paper.

Probing Visual Language Priors in VLMs Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T22:55:54.242012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:55:54.242012Z digest=sha256:e873b2acbb21c2dd6301e8329102e0cda93a2cb4abc72506d0d5497b67f0737f

Observation 5b90f82b-6a40-4934-81b6-1990ccbb9a7c · inbound

A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges cites this paper.

A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 224

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:23.786392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:17:23.786392Z digest=sha256:9136cc5a315a04425c5815d288bb3d1be6a153620cf4be1c2bc6dc100cbd87a6

Observation 407d1acc-13bc-4ac8-89f7-d4b7fceafdf1 · inbound

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision cites this paper.

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T21:34:35.564572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:34:35.564572Z digest=sha256:8916c4f72af8b68d553255238b96cde354e65cd167e73826d589cd67d1957df0

Observation e6eb6a9e-81eb-4c95-b207-e00b123844d9 · inbound

MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs cites this paper.

MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T17:01:41.205261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:01:41.205261Z digest=sha256:90c5cfbc3c70e9f13e782fba4c59ce7877524b515deb72e473eca9374b25185f

Observation db392c99-de85-4b05-bf75-ade15cb09750 · inbound

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning cites this paper.

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:59.558385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:22:59.558385Z digest=sha256:2e31d0deca4f000f0a987e3f3260232c776ac5c75f52d3e60bc5ec050e0f41f4

Observation e873ec0a-ab0b-4f29-a3b5-edc9fc45642c · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:52.987620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:52.987620Z digest=sha256:b2789b134e65b06239563a3fcd25f473a3b1d47f2685d368132ed332362f3ac5

Observation 677d734b-80d1-4bf2-b4ce-31d3dbac3d35 · inbound

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs cites this paper.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.660360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.660360Z digest=sha256:f9464069a3634d725df39cc0bb519d262cf120162c0a72cf4ae368cdbc36ffc8

Observation 5f9f30fe-f43a-4384-b538-f6b6581fae99 · inbound

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning cites this paper.

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T18:39:18.941061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:39:18.941061Z digest=sha256:728994d6260c5fef0d10371a81ce3b0a6d82079b9ec5487c0e4377d136b1acf7

Observation 898c1883-d398-4829-a80a-277d07070c79 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:36.307866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:36.307866Z digest=sha256:5de9e09b2f0df514d9f1122cc5222dcdfc3adeba61b16ca96a0f8eb49d53c87e

Observation e080abb2-18c3-479e-9589-b312c8f11082 · inbound

Chart-CoCa: Self-Improving Chart Understanding of Vision LMs via Code-Driven Synthesis and Candidate-Conditioned Answering cites this paper.

Chart-CoCa: Self-Improving Chart Understanding of Vision LMs via Code-Driven Synthesis and Candidate-Conditioned Answering Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:32:08.998266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:32:08.998266Z digest=sha256:59d82dd6d2121efd2ede86f5af24e343437992fa7954961d9a590b6061ae95f7

Observation 05e5b6d7-ed0c-4f79-b7bb-65ec5234e508 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.852228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.852228Z digest=sha256:02c3245901ee652a17cf036792d29416190a4044de706c904b11da059ae1a205

Observation fe9fa07f-4118-41da-b60a-83beba9e54c7 · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.628450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.628450Z digest=sha256:c3b7943773967659b076df5e331b40b86016e633ba42b006f05b29a8b682bb0c

Observation 89b73dba-f893-46b0-9698-6d8270c23677 · inbound

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? cites this paper.

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:36:18.594594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T03:35:39.884350Z digest=sha256:eac4e496e3957907367e1e43f89b442fb7035dec5bc57a3fdbc6c39ee7915074