Pith. sign in

Paper Citation Record · LEDGER

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning

As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2506.15649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15649 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:57:22.557555Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:38:51.797197Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T08:26:00.654417Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3a0d8e4-1821-4a38-9110-cf15f90b77af · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:19.314620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:19.314620Z digest=sha256:b14abdc4f7990d1d78a5afaf82d09be8b34b5f61e5d504b6f40d21618e528c3a

Observation 4387977e-f43e-4211-8819-1a8435f62ed2 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:19.351208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:19.351208Z digest=sha256:f7481f6e612f6f46603c0fe33c29f931e107225e4137bf5461e61dc54bce480e

Observation 97e5e07f-fcfb-49e7-a6c9-e5a55985f41c · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Hallucination of Multimodal Large Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:19.398994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:19.398994Z digest=sha256:812e84ccc8f4cef91944c26ea6158551837ec886bc005e273af5c113fe92ec3a

Observation 5f8edd4f-aead-46bc-8f9d-25f3de7ae660 · outbound

This paper cites Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:57:23.010234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:19.479717Z digest=sha256:c22ef32c2dee26db0d3154da3b76160b22319fd62568b9b5c16c9f7515fe65a2

Observation ea2a9ef5-92ed-4008-9584-8529a20b3389 · outbound

This paper cites Improving image generation with better captions.Computer Science.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Improving image generation with better captions.Computer Science

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:19.569996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:19.569996Z digest=sha256:6233d9b3aaa4922eb4fa835df9259f69a8ff9c4aafcbbbc307e4f1b3c1102962

Observation c8f90c75-aef8-4146-8e3f-097f8e16b0da · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning PaliGemma: A versatile 3B VLM for transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:19.619534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:19.619534Z digest=sha256:7966c10663aa823fe88314f00fac26bb41dc58d87e7c93512cfa442018f4c0fb

Observation ffb6c982-6d86-49ed-b1be-74d1a4f825fd · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:19.701437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:19.701437Z digest=sha256:d0eae8fadfd88fb24c7747a786ed92fae2c77e7bdee83dff44f1a4b49758cbae

Observation 32b47546-e20b-42a8-82f6-4ea6ca981cfb · outbound

This paper cites Transfer q-star: Principled decoding for llm alignment.Advances in Neural Information Processing Systems, 37:101725–101761, 2024.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Transfer q-star: Principled decoding for llm alignment.Advances in Neural Information Processing Systems, 37:101725–101761, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:57:24.240903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:19.773228Z digest=sha256:400de4b42e40dacf92346bb41fe45dc4e6202c3e0dce3e2e1d14347966d19f47

Observation c7a7959a-32c5-4ef5-893f-980207911ae5 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Sharegpt4v: Improving large multi-modal models with better captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:19.841908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:19.841908Z digest=sha256:0918030d07983cb042109b1f35388c8673cd26e639efcab2412189a886156c2c

Observation a16795af-8bba-4a70-91b0-976e0e609a9d · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:19.883212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:19.883212Z digest=sha256:9e96f1983d2b7acaadf3a044153e17e8eba6826ecd23af46bf7b471d37da5583

Observation b5109b54-d3fc-4f05-8b12-37b10a87a8aa · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:19.973475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:19.973475Z digest=sha256:94c7b94a105e946f14801edffb5ee6b88e543613ed088126498f74ccc6822ccb

Observation a8ac9528-9e73-42c8-9331-30088b03dc91 · outbound

This paper cites Mitigating Hallucination in Visual Language Models with Visual Supervision.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mitigating Hallucination in Visual Language Models with Visual Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.052718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.052718Z digest=sha256:3dafcda12dec620620ed3b10e406010a02dba010dc314b38ba7dcdf3858d0494

Observation 312f131d-840f-4bf8-91b6-405b875d5e18 · outbound

This paper cites Partially non-autoregressive image captioning.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Partially non-autoregressive image captioning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:57:24.146107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:20.088656Z digest=sha256:3aa86603844789eeca8e1b72789deda22cf5467a882cbd0eb4235b645e037bad

Observation badfbf07-825b-46b9-ab67-e9aa76ea2750 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.164694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.164694Z digest=sha256:4ca0c4cb9988300d664b5a6fd59d22270b9c4eed9875112a1666ec3d51465afe

Observation 2fd181c1-01f0-4488-b6ba-9ebc6bc22ed1 · outbound

This paper cites From images to textual prompts: Zero-shot visual question answering with frozen large language models.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning From images to textual prompts: Zero-shot visual question answering with frozen large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:57:24.026884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:20.230133Z digest=sha256:12f94ef6988665b269441e8778f7dcfe7e6847ce34086026831c9acb398ad1f8

Observation baa8ed80-8e2f-4d1d-a509-02891df79859 · outbound

This paper cites A hierarchical approach for generating descriptive image paragraphs.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning A hierarchical approach for generating descriptive image paragraphs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:57:23.943653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:20.267102Z digest=sha256:cc6e6c0daa3bd171d9b9e7414efe90ade55b2decdc02b6fd705d20ec8b11cb0c

Observation b8aad40e-88c7-4639-bb20-8120b7c18a1f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.342916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.342916Z digest=sha256:b0cdc9b4bdc6428bd71771523c9a27bcc9da9b8d6c348926b8b39a87c7322171

Observation 9799f53a-05cd-4ff7-ba67-53558097a6cf · outbound

This paper cites Coco-cn for cross-lingual image tagging, captioning, and retrieval.IEEE Transactions on Multimedia, 21(9):2347–2360, 2019.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Coco-cn for cross-lingual image tagging, captioning, and retrieval.IEEE Transactions on Multimedia, 21(9):2347–2360, 2019

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:57:23.836349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:20.429119Z digest=sha256:f93e4a227621a5c4b7ba0846a9485ca8f9027ec0759baa7daa6d06ba3d4a8484

Observation fff19fa4-3208-4752-b9a2-b4b0568f9195 · outbound

This paper cites Let’s verify step by step.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Let’s verify step by step

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.468101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.468101Z digest=sha256:02f45f0abd20e9098c65a9ddb1efe9170fcd81f94a801b26db31ec2baed81968

Observation 0ada43e7-ed5c-4ea4-b2f9-f633f5b67dea · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.520182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.520182Z digest=sha256:e021b1ced38acae8dffef2985273731d0ee74114c3720b3ac5fe170315d2addb

Observation 7f11a5f8-2f89-42b1-ab94-beb7b8a371f1 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.613758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.613758Z digest=sha256:bcd6ff0f88a0acf63d504555f535445748ff25f935d9106acb1d12ec8b5511b2

Observation 91f2a0fe-604b-432c-bcb3-1d95858b90b1 · outbound

This paper cites Improved baselines with visual instruction tuning.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Improved baselines with visual instruction tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.681130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.681130Z digest=sha256:c16ec3854e76132b0936f93a7e41c3de7397a23d350825c2dc5b28a62cb72b4d

Observation 18677114-1d80-410a-a179-95fa8b187deb · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.729906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.729906Z digest=sha256:e7a4f12d6e6c8958138ffc90c1afacb892de1e04d8d00b4e405179ee6f53a328

Observation 25a2a333-4f8a-4cc7-b44f-1d2f49254ae2 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.762953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.762953Z digest=sha256:b5bf6af8fdbc09c8782c14b881280f5b2382e1675ff89d93fb716c1ca161526b

Observation 4604e7ca-aa56-4c33-8770-07fa0d09d579 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.836219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.836219Z digest=sha256:0813a0cf1decd7a31a669ec7ba873fa8ab8bc4c0a142b612c555bbc67ed0a679

Observation bada4063-79d4-4fb4-8989-5b7a52968bdd · outbound

This paper cites Hallucination Detection and Hallucination Mitigation: An Investigation.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Hallucination Detection and Hallucination Mitigation: An Investigation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:20.923852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:20.923852Z digest=sha256:b603fb0f6e6e3370a68f440f86ec8703a35d1ffd16d01999addb296c92583983

Observation 9f9aa8c8-0eec-4fef-8209-be6c9066edff · outbound

This paper cites Training for diversity in image paragraph captioning.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Training for diversity in image paragraph captioning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:57:23.668793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:20.960011Z digest=sha256:0eeb992ad7f21572efe1d10fe314ac018c6dc1a8a162e6f20d1ab1ffaece2093

Observation c971c67b-3347-43c6-ade0-6c746d2c4b06 · outbound

This paper cites Learning to reason with llms.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Learning to reason with llms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:57:23.557590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:21.029431Z digest=sha256:6fa0cea0b9e456e2479f8555e9bc60bf77ae501865fb78e13e2b3155f1d64b82

Observation e4cbf877-9d8f-4e75-b4b5-059dfa50bd70 · outbound

This paper cites Self-critical sequence training for image captioning.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Self-critical sequence training for image captioning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:57:23.445809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:21.106679Z digest=sha256:3bcb2d27cf1beb46ea5d0fdd06aa0e904208bca98c8c6b935f10c88c8d09809d

Observation 01b005c7-93ab-47be-a9ff-667ac89ae176 · outbound

This paper cites Object Hallucination in Image Captioning.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Object Hallucination in Image Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.157328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.157328Z digest=sha256:b82b16230775be4ec10a7834b8f9f05fcef36ffb6183660329fd8b0ae41b89c7

Observation 9f52c8c8-8e9a-42a8-ba63-7d0a5387edbb · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:57:23.291456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:21.186831Z digest=sha256:c81458fb53d2ddfb4b690bbe9c899f8d44608c6db2d0f3eeb0aa32f7dc2d89eb

Observation 22104888-65fe-4ced-bb63-bfa195d2538e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.272982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.272982Z digest=sha256:e77ad9c99d4edac0704b1821a72167609f2aaf1c55c1a5ca6919e234f555cd92

Observation 2083ec24-9664-4998-aa1a-381b6a7f2f86 · outbound

This paper cites The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.346054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.346054Z digest=sha256:85ec18d17a33f25d554d0f761f186a5d39f118561a7f4bcea799a4ff096aedb6

Observation 2badc4ec-3dd1-4000-acef-6cc333f5ca9b · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.395412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.395412Z digest=sha256:a8f7e77dff54a48ec8ab885bda8276064807850fae9b37457d9bfbf581d1708a

Observation 4a53ce7c-c36b-49f2-a117-8438ed88c61c · outbound

This paper cites Learning to predict by the methods of temporal differences.Machine learning, 3:9–44, 1988.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Learning to predict by the methods of temporal differences.Machine learning, 3:9–44, 1988

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:57:23.199958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:57:21.483237Z digest=sha256:2c6c3e830186a05577d1ee55f8b34dd4eaf56c3d5983a185584df97957f38c66

Observation 6161da9a-0e05-440f-8652-c9dc3507250c · outbound

This paper cites Toward self-improvement of llms via imagination, searching, and criticizing.Advances in Neural Information Processing Systems, 37:52723–52748, 2024.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Toward self-improvement of llms via imagination, searching, and criticizing.Advances in Neural Information Processing Systems, 37:52723–52748, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.554884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.554884Z digest=sha256:86b8f026e8c0f119772fe8ca92db0d9bc7b84bc9eec4d84c406c0371926992c8

Observation eed26693-5325-4bc2-8f2d-dfd8fae7300d · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.596476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.596476Z digest=sha256:002882429a4be8bf728447fc0ce160942cabaabf95ba86b166226e567426818d

Observation 958661bd-7dfd-4d18-b87e-2c3dfbd1adad · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Solving math word problems with process- and outcome-based feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.698625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.698625Z digest=sha256:4c3b319920f230c1bf432e364677d2483490a56bb5237e712eaa0242689f10e3

Observation 6b86856d-6d8d-41c2-b1f6-810e68495616 · outbound

This paper cites Faithfulness-Aware Decoding Strategies for Abstractive Summarization.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Faithfulness-Aware Decoding Strategies for Abstractive Summarization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.762431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.762431Z digest=sha256:37a8d2dd5a7b81b365772739f9d6c4ecfda885ee8fe8f949a6381f4408c1a3ac

Observation c3f05ff3-ff21-47a7-97b6-7bcc71065d9e · outbound

This paper cites LiteSearch: Efficacious Tree Search for LLM.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning LiteSearch: Efficacious Tree Search for LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.815062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.815062Z digest=sha256:2f31a7d998daba70bb99d2d1d810cf79fd7f3bb78f10bbd9c7e23b0184c0f940

Observation ced584f3-ad29-4fc9-98b9-c08ac9daf041 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.877950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.877950Z digest=sha256:9ccdd8e84c2495c905eaf84f3ac9e3c7f1203e34fb05ce898bff2c323f3d0064

Observation d31f81cc-df1b-4f30-9124-8d52316a7691 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.Advances in Neural Information Processing Systems, 37:121475–121499, 2024.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Cogvlm: Visual expert for pretrained language models.Advances in Neural Information Processing Systems, 37:121475–121499, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.930665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.930665Z digest=sha256:6c7945193a97b3041201d411fab6c24dca1b9839f4b56e6a633f83d15eac98b0

Observation 4b54e66f-c402-4f47-a302-c723e16ab620 · outbound

This paper cites Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.985073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.985073Z digest=sha256:748e2a396b2eedffa0e08fe2df1892792dd1424b0872fcd597a0bbec1d5aac6d

Observation 6add957d-08f8-48a0-9bdf-d3f310ccc30a · outbound

This paper cites Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.096140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.096140Z digest=sha256:5cc939709ef41aee0455b58406bf361845dc8ecc75355b2979433766dc11c342

Observation 7b28090b-ee75-4c24-8f16-201a27e47f49 · outbound

This paper cites Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.137629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.137629Z digest=sha256:4f0282a0023763723da9a52a2d591e12b1d9a1b4afacc97514759d7378abc8bb

Observation da741311-9950-4199-9b67-7ce3d572503b · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.189577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.189577Z digest=sha256:bca95964c83bcabb2ca309fac5ae72a0bc51bbd8c5f3994694452517f8868398

Observation d599e882-1779-426b-a9c6-f0a1c124310d · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.277104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.277104Z digest=sha256:46443fa0fa7f23e0122a4124d5ee929299f6b86d8f63919ee95653f87dbcd1a7

Observation 8a27f66c-70d5-462d-807b-3844fd4a638c · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.318324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.318324Z digest=sha256:809f2e12c69b0b7dc1231457d132b8521d0cfd31002dfb461963e942cb2dadd3

Observation b11da14d-8979-476d-9afb-88e9fc4db042 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.370183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.370183Z digest=sha256:008ef19491abb222646adeffbf0eebb739442601965bedbea90d98ede884d084

Observation d656ce03-b75d-4f32-9ef3-e11324724b83 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.453247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.453247Z digest=sha256:5ce55ab1000cedc7e08a102469b96e3a1c0e2fcb7ec021ddef9413f6138fee49

Observation 6a7d0348-bef1-47bd-aac9-8be53d539b74 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.498077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.498077Z digest=sha256:31dd45e43786c704b750c2ffd2c057c441b06fdd531ebac48ff38b4d4492fc4e

Observation ba057781-2d4e-43f2-ad73-e77052e5b73c · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search.Advances in Neural Information Processing Systems, 37:64735–64772, 2024.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Rest-mcts*: Llm self-training via process reward guided tree search.Advances in Neural Information Processing Systems, 37:64735–64772, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.527750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.527750Z digest=sha256:4c66089e5fc4b75b993616dee9b02d0dc657678093cbc51699a2d6a4ba069155

Observation d78dd73f-5a7f-462e-bac7-f4792db5b05d · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Calibrated Self-Rewarding Vision Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.557555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.557555Z digest=sha256:661d82036e1fbecf2e7a3f0e5846f0fa2c9b2595a8f62052fc4a988f56f3464e

Pith citing papers

Observation e0d91c50-46db-4b40-97dc-09cd1570a794 · inbound

See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment cites this paper.

See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:26:00.656565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:38:51.797197Z digest=sha256:8beac1c654b287f5126c99fb0aa19cbdff9dbf0ef03d4d11f68ab824e9390a7d