Pith. sign in

Paper Citation Record · LEDGER

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2412.03704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03704 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:12:06.014106Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T13:24:40.386203Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b69cd9cc-54b3-45bf-9283-819838cacd21 · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:06.014106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:06.014106Z digest=sha256:51295cef44aef1ed1e9ed3ba4831762e241c489eba5594b73abddae5a90e8337

Observation 246c9b34-c7f0-4ca4-bee2-4043cdf08d90 · inbound

Mitigating Object Hallucination via Robust Local Perception Search cites this paper.

Mitigating Object Hallucination via Robust Local Perception Search Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:03.181743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:03.181743Z digest=sha256:7fbbc5a8d60d63140f42b1b05cd30972b0408a8fd23cfbd2e8d28cfbba807b43

Observation b54b5691-eafa-49dc-87b5-550f54f6c437 · inbound

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding cites this paper.

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:18.661384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:18.661384Z digest=sha256:eef2156a120a6253c1b9fc010ebb2412701705d89a729f965694f69ab52cc4ef

Observation d0da67b2-01d5-4d05-a158-b71c70f17dfa · inbound

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs cites this paper.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.741673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.741673Z digest=sha256:9bf35c22838532a1d8fe461d9a1fd6bfa75861cf629f576d26ee8b661d9dd658

Observation 6add957d-08f8-48a0-9bdf-d3f310ccc30a · inbound

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning cites this paper.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.096140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.096140Z digest=sha256:5cc939709ef41aee0455b58406bf361845dc8ecc75355b2979433766dc11c342

Observation e74bdea1-5859-4b61-a8b2-633f4646b1b9 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:24:40.393917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.858355Z digest=sha256:1c320d534a60d79db9f2c93c3e0532f8bd31b1099b93d1130baeaa1d1f83431f