Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:57:22.557555Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2506.15649.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:57:22.557555Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:38:51.797197Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T08:26:00.654417Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b3a0d8e4-1821-4a38-9110-cf15f90b77af · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4387977e-f43e-4211-8819-1a8435f62ed2 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e5e07f-fcfb-49e7-a6c9-e5a55985f41c · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Hallucination of Multimodal Large Language Models: A Survey
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f8edd4f-aead-46bc-8f9d-25f3de7ae660 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea2a9ef5-92ed-4008-9584-8529a20b3389 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Improving image generation with better captions.Computer Science
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f90c75-aef8-4146-8e3f-097f8e16b0da · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning PaliGemma: A versatile 3B VLM for transfer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb6c982-6d86-49ed-b1be-74d1a4f825fd · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32b47546-e20b-42a8-82f6-4ea6ca981cfb · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Transfer q-star: Principled decoding for llm alignment.Advances in Neural Information Processing Systems, 37:101725–101761, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7a7959a-32c5-4ef5-893f-980207911ae5 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Sharegpt4v: Improving large multi-modal models with better captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a16795af-8bba-4a70-91b0-976e0e609a9d · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5109b54-d3fc-4f05-8b12-37b10a87a8aa · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ac9528-9e73-42c8-9331-30088b03dc91 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mitigating Hallucination in Visual Language Models with Visual Supervision
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312f131d-840f-4bf8-91b6-405b875d5e18 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Partially non-autoregressive image captioning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation badfbf07-825b-46b9-ab67-e9aa76ea2750 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd181c1-01f0-4488-b6ba-9ebc6bc22ed1 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning From images to textual prompts: Zero-shot visual question answering with frozen large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation baa8ed80-8e2f-4d1d-a509-02891df79859 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning A hierarchical approach for generating descriptive image paragraphs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8aad40e-88c7-4639-bb20-8120b7c18a1f · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9799f53a-05cd-4ff7-ba67-53558097a6cf · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Coco-cn for cross-lingual image tagging, captioning, and retrieval.IEEE Transactions on Multimedia, 21(9):2347–2360, 2019
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fff19fa4-3208-4752-b9a2-b4b0568f9195 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Let’s verify step by step
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ada43e7-ed5c-4ea4-b2f9-f633f5b67dea · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f11a5f8-2f89-42b1-ab94-beb7b8a371f1 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f2a0fe-604b-432c-bcb3-1d95858b90b1 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Improved baselines with visual instruction tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18677114-1d80-410a-a179-95fa8b187deb · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a2a333-4f8a-4cc7-b44f-1d2f49254ae2 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4604e7ca-aa56-4c33-8770-07fa0d09d579 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bada4063-79d4-4fb4-8989-5b7a52968bdd · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Hallucination Detection and Hallucination Mitigation: An Investigation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9aa8c8-0eec-4fef-8209-be6c9066edff · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Training for diversity in image paragraph captioning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c971c67b-3347-43c6-ade0-6c746d2c4b06 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Learning to reason with llms
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4cbf877-9d8f-4e75-b4b5-059dfa50bd70 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Self-critical sequence training for image captioning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01b005c7-93ab-47be-a9ff-667ac89ae176 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Object Hallucination in Image Captioning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f52c8c8-8e9a-42a8-ba63-7d0a5387edbb · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22104888-65fe-4ced-bb63-bfa195d2538e · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2083ec24-9664-4998-aa1a-381b6a7f2f86 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2badc4ec-3dd1-4000-acef-6cc333f5ca9b · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a53ce7c-c36b-49f2-a117-8438ed88c61c · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Learning to predict by the methods of temporal differences.Machine learning, 3:9–44, 1988
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6161da9a-0e05-440f-8652-c9dc3507250c · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Toward self-improvement of llms via imagination, searching, and criticizing.Advances in Neural Information Processing Systems, 37:52723–52748, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed26693-5325-4bc2-8f2d-dfd8fae7300d · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 958661bd-7dfd-4d18-b87e-2c3dfbd1adad · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Solving math word problems with process- and outcome-based feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b86856d-6d8d-41c2-b1f6-810e68495616 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Faithfulness-Aware Decoding Strategies for Abstractive Summarization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f05ff3-ff21-47a7-97b6-7bcc71065d9e · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning LiteSearch: Efficacious Tree Search for LLM
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced584f3-ad29-4fc9-98b9-c08ac9daf041 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d31f81cc-df1b-4f30-9124-8d52316a7691 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Cogvlm: Visual expert for pretrained language models.Advances in Neural Information Processing Systems, 37:121475–121499, 2024
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b54e66f-c402-4f47-a302-c723e16ab620 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6add957d-08f8-48a0-9bdf-d3f310ccc30a · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b28090b-ee75-4c24-8f16-201a27e47f49 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da741311-9950-4199-9b67-7ce3d572503b · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d599e882-1779-426b-a9c6-f0a1c124310d · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning LLaVA-Critic: Learning to Evaluate Multimodal Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a27f66c-70d5-462d-807b-3844fd4a638c · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b11da14d-8979-476d-9afb-88e9fc4db042 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d656ce03-b75d-4f32-9ef3-e11324724b83 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a7d0348-bef1-47bd-aac9-8be53d539b74 · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba057781-2d4e-43f2-ad73-e77052e5b73c · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Rest-mcts*: Llm self-training via process reward guided tree search.Advances in Neural Information Processing Systems, 37:64735–64772, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d78dd73f-5a7f-462e-bac7-f4792db5b05d · outbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Calibrated Self-Rewarding Vision Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d91c50-46db-4b40-97dc-09cd1570a794 · inbound
See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.