Pith. sign in

Paper Citation Record · LEDGER

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

As of 13 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 8 inbound Pith citation observations for arXiv:2412.03704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03704 v3

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:18:01.740237Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:18:41.089878Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T13:24:40.386203Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved45
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8e51b86-0b81-4d14-805d-e682e575598f · outbound

This paper cites https://openai.com/ index/learning-to-reason-with-llms/ , 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension https://openai.com/ index/learning-to-reason-with-llms/ , 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.574806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.374966Z digest=sha256:6dd0ef3731eaf0203036cdf432e535ba7b5bdb1348b145b6c8968cd36365398b

Observation 9a0e4367-0052-46e6-bb5f-fa2511edaac3 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.380204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.380204Z digest=sha256:c57728be1d5fa4ae77ca72472bb61433f668dd7eeda9a8d5d9d3589ec21dae12

Observation 7e85f58e-88a0-4cb7-ab22-cfc79e54d146 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.384867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.384867Z digest=sha256:971958f739ae1c5570ff645232cc2fddbb3dd4a77450af3dc032f8fea69fd1ec

Observation 0b621301-cb13-4419-a8c1-5bb2186a4602 · outbound

This paper cites Improving image generation with better captions.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Improving image generation with better captions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.536833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.389776Z digest=sha256:2256f9992e0e1151f559b743aa249eea73f8f4d5164e3667c7d2fc250a7f3e7b

Observation 2c120969-fdb5-4970-893e-4efbcaca9674 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.394336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.394336Z digest=sha256:6a6ffbef1f13ace39233ace73cd23aa0c5ddcdf9b7dd4985797a9e560947ca1e

Observation 06c1ee18-8307-4e8c-9f93-ada6a01830dc · outbound

This paper cites Transfer Q Star: Principled Decoding for LLM Alignment.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Transfer Q Star: Principled Decoding for LLM Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.399147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.399147Z digest=sha256:5d1d0aca233f611f6409853576c6402bbad50272d2e8eb6d09bb091a45a7f7bd

Observation 64a4b06c-2c93-4cdf-a64e-07dcfb3bd8ea · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.507090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.403556Z digest=sha256:b8c330f5a638a18b768f9039d231f21ab9ae29b1af1c7be7942195f84e9e276e

Observation f8594860-4568-4fcd-98a5-7b1f76e26593 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.408254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.408254Z digest=sha256:d5653de7aeed8fd9c78e06eb9c45a8bf01449e6a98a4ad2abdc8c3bb79149cbc

Observation a4521f82-edf5-43e2-a32f-ca88a8b9f033 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Are we on the right way for evaluating large vision-language models?, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.473365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.413329Z digest=sha256:c591d939946fa2793687b0ae5553b044ac824b37927e09707fc147a84590437e

Observation afe682d7-db3e-4580-ab11-ef480f6352dc · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.417584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.417584Z digest=sha256:0ae5961c07d6641499b9eecd3452b3810e4ab04673840f3e309f3807eda49c72

Observation 681c3297-6802-4805-b162-e95bff58f360 · outbound

This paper cites Mitigating hallucination in visual language models with visual supervi- sion, 2023.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mitigating hallucination in visual language models with visual supervi- sion, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.442540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.422359Z digest=sha256:e221a22854dec113f68024dcb65159c66e5bd03d8537235b066a89dc6564e949

Observation a353d46b-873a-4260-9dab-6cd45a1482e0 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.426734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.426734Z digest=sha256:61a8c7ebc4f788e111746bd4855b19cc6a2e4d5a6d645a697c533957c4c34cf2

Observation 2ba321c9-89fa-448a-8b85-73b7439afbd0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Training Verifiers to Solve Math Word Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.431073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.431073Z digest=sha256:3ff638e15ab9986ccda0fd954a85166766562461523c67cc9706d8fc008e0da0

Observation 9e41440e-4856-4ea2-ad43-5bba36692dc1 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.435361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.435361Z digest=sha256:f55b590e631b8d1b730f97708b42fd7e9eb785825e2a218a091e023751560fd7

Observation d2db04c9-28c5-47ae-b547-e07b993b08e3 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.439358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.439358Z digest=sha256:b00a2418dfbc1531a184f1461f5b37b1381b858b59466d34d717fb7a88edf816

Observation c1659d40-3742-42ec-97fb-5f673502d2e9 · outbound

This paper cites Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.374022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.443815Z digest=sha256:20e367c0c87255cd234d293752398405687945c561eab5f41ae233ef9f01a2df

Observation fd60810a-6161-49f1-9ab1-3f5afd1aad28 · outbound

This paper cites Temporal difference learning for model predictive control, 2022.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Temporal difference learning for model predictive control, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.350810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.448666Z digest=sha256:dc392f0e2020dc6df03ab6c049a432d075c6e38f00bac9cb83839d3e89cae8dd

Observation 12ddf2a6-6e41-4fd8-a2dc-360e4e7cca1a · outbound

This paper cites V- star: Training verifiers for self-taught reasoners, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension V- star: Training verifiers for self-taught reasoners, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.330783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.452970Z digest=sha256:fb9fd747b340708d72228fabfb0c82585a9dbdd7a7f1c30d584a04de74f765f4

Observation 3f1690e4-1b81-4e25-9fa5-90509b98b4a6 · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling up vision-language pre-training for image captioning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.310213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.457983Z digest=sha256:3462ea31e02b313e6c7a5753cd85784e20a1ec88cae0aa0e9c91ca1c5f91205a

Observation 6e4539fe-dfb6-4e89-bea0-28536504ee53 · outbound

This paper cites Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.462379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.462379Z digest=sha256:7bf9bdcf7fd1381dd650131607a14f364787482539cd76ff34763da950223479

Observation 916b628c-fa8a-4929-a6ce-f7054f106ace · outbound

This paper cites Mantis: Interleaved multi-image instruction tuning, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mantis: Interleaved multi-image instruction tuning, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.286821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.466985Z digest=sha256:c956bed02011c67e999732b9f81e4b8e56f43a73ad7eb3d4e6233faac0eb9395

Observation 01a791ac-b095-4a8e-b28d-7419fd44d011 · outbound

This paper cites Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.471331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.471331Z digest=sha256:d0c6285a3984cb0b4d786b138b5b9ca44097ab38f93c350b16d6cde5be7d76cd

Observation 69012f65-90d3-48f5-b595-e569857746ea · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Veclip: Improving clip training via visual-enriched captions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.267991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.475668Z digest=sha256:8adc1b1210459763e32e56b37e8e610f19357631a501e09ae50e070d38a446cf

Observation e30b6298-7524-4f23-887e-acef171dc756 · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding, 2023.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.247813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.479955Z digest=sha256:8e47c58a7881c04a31e8e21348dcc3d386762045cd0cdc63599dac7229d1810e

Observation e73729a0-9c58-406a-b897-4790b967c806 · outbound

This paper cites Llava-onevision: Easy visual task transfer, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Llava-onevision: Easy visual task transfer, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.485065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.485065Z digest=sha256:7a5602f8c6543a6ba9799eaf868d2ae3ffd75a3c4f30f5f87676be85657f3183

Observation 8d36b022-486b-43cf-a045-d18933759d93 · outbound

This paper cites Multimodal foundation models: From specialists to general-purpose as- sistants.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Multimodal foundation models: From specialists to general-purpose as- sistants

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.219221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.489185Z digest=sha256:b4cb6e7de440cff9647661b1ae8c49bc9b61268be63fcf9681e009b9a2a8c2d0

Observation 0d2c09b7-05d3-455e-89dc-c39153e4bf1d · outbound

This paper cites Common 7b language models already possess strong math capabilities,.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Common 7b language models already possess strong math capabilities,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.200570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.493295Z digest=sha256:19abc37c918975aa57e559a278d7b6b77c21bcaf2b1c5c976d030897127c1b23

Observation be12431c-c9e1-454b-89b5-b82695d36f8a · outbound

This paper cites Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.497490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.497490Z digest=sha256:5574c3e75eeb98897f3f661a5127ecb07a4616138f6becf5397e85ded0014d0e

Observation 39a14e9e-12e9-4a08-832c-9b279b9204ea · outbound

This paper cites Let's Verify Step by Step.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Let's Verify Step by Step

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.501651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.501651Z digest=sha256:853918a0ec81bd079c22424db9197d3596dd05c86a51ad54ecf2b698bb8b2df6

Observation 26e4f717-0b73-43bb-b762-6220ff9e8d66 · outbound

This paper cites Let’s verify step by step,.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Let’s verify step by step,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.506210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.506210Z digest=sha256:ebf13ab42239748e7321fa39a74ecf8865a14bb81abcb80982aa045d0c226044

Observation 8814eb9d-d90f-4304-b8bb-3e3b10fdb5b0 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.511015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.511015Z digest=sha256:7dcffec84c78d5ac1d3b778b48e57e5fd8061ecc5319ff3364d54547f050e73e

Observation 36b18720-377f-4bbf-87cf-2c7631b8151e · outbound

This paper cites Visual instruction tuning, 2023.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Visual instruction tuning, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.516148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.516148Z digest=sha256:32ee590b519f33450ec10b3a9fe9092388e46e07961f7e6f5e631d4c37a30804

Observation 19793b71-9131-4b85-bf01-21154a796e88 · outbound

This paper cites Visual Instruction Tuning.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.520496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.520496Z digest=sha256:8788c49dfe9f21b23245557060c356d5639c9d84f1de55db25e81c4bd6c10d9a

Observation e772ed4e-6e5a-4d0e-ba61-74d8bbab8b85 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.526266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.526266Z digest=sha256:344a52961d5dccbd59139e49e4c4c2f145572faf9edddb72b2bff3282062f283

Observation 23975e21-4eae-4fa0-898f-8c7b2c7e4fa8 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.531017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.531017Z digest=sha256:a0d188fbfe6c81322e5c70d80877da6f1e63bd4b51e7296c7102777f708dbaf6

Observation 2121e488-736c-41bd-93d8-8694a3c22ac2 · outbound

This paper cites Gpt-4v(ision) system card.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Gpt-4v(ision) system card

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.535511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.535511Z digest=sha256:328b5bf655c6984d1ea20701974000bccfc4b159cb5350535b3987d417e14cc5

Observation 81191c14-d97b-4a90-857f-8d9bb695ccee · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Learning Transferable Visual Models From Natural Language Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.540118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.540118Z digest=sha256:cb9e6c8d05c710f2acab8fb5f477da3fbe676bd0c1f37b194839ba5975ce858c

Observation 1bbe461c-859e-45e6-800c-2fe59be2c1a0 · outbound

This paper cites Object Hallucination in Image Captioning.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Object Hallucination in Image Captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.546014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.546014Z digest=sha256:497a248a596e2e00d3882f9c10a26534f94db91c87e9b47782301a528b63aa28

Observation 16fb8192-cc40-498d-9fdf-55ca80f0de8c · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.551747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.551747Z digest=sha256:ad388e5f117dd6e50122a38b749d6670bebaa5631c829eaa34c269cdd8476087

Observation 7308baae-38cc-4f4b-afe7-a3bbaceb799e · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mastering the game of go with deep neural networks and tree search

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.094201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.556427Z digest=sha256:48739199689432f7eb850c1e62e8f90a6ccf66a126714b2da5a69ee4f1eecacd

Observation 74df077d-7cc5-4228-b980-3674728ecef1 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.561998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.561998Z digest=sha256:a2e0694d9d97d905b31c84fc8521a4f0c9559d580a50c5c8a68737c3dfc38e8e

Observation 29867786-e81b-48c7-839c-bc0b6d256b49 · outbound

This paper cites Scaling llm test-time compute optimally can be more effec- tive than scaling model parameters, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling llm test-time compute optimally can be more effec- tive than scaling model parameters, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.568322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.568322Z digest=sha256:18315c95efed9bc2c2e8bf02df3ff6b453113c2f0ae07b66dd38961d9d1ad213

Observation 96572944-b38b-4991-9271-9f1a02ec53c0 · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf, 2023.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Aligning large multimodal models with factually augmented rlhf, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.058215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.575285Z digest=sha256:ac842025293978b928b65de9544f7aa7b9ff4c96a7ccf88c80fb720d6ec7eec5

Observation c167b5dc-8969-4f95-a088-0677866ad949 · outbound

This paper cites Learning to predict by the methods of temporal differences.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Learning to predict by the methods of temporal differences

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.034643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.580916Z digest=sha256:d9dc3d1a1cc16cf59bfb3590c1962c0c711bbd31a09ce148304d7bd157dc0c10

Observation 0e8820ad-e1b2-49c7-893f-85e1f01cd99f · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2023.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Gemini: A family of highly capable multimodal models, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:03.009850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.586993Z digest=sha256:4d2d13c8a279b52bf6f0fbf274add32e70ec6b2b36881004a2e4acc5e1ec91a2

Observation 26273c98-836d-4bfe-9810-fca71c780a5e · outbound

This paper cites Motion planning for au- tonomous driving: The state of the art and future per- spectives.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Motion planning for au- tonomous driving: The state of the art and future per- spectives

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:02.985041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.592542Z digest=sha256:dfe3914934764da01313d3d4307d97e6ca390c54f3442646e9071a4de0dab475

Observation b40c4073-d9d6-41ce-b7dc-b47703ac6c4f · outbound

This paper cites Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.597210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.597210Z digest=sha256:c20e024504665eaf7a3a1d135f4b7103f4e99e46d5fa58bdccebba95908e1c6a

Observation afad285e-165d-49da-8831-af19fe27bde2 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms,.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Cambrian-1: A fully open, vision-centric exploration of multimodal llms,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.602606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.602606Z digest=sha256:0ae208993664d40239bedb4a0ab1f3929cec7918539936f3f2a972cc8672a9cc

Observation 22ab22c4-fa45-4098-96f7-b799195064fa · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Solving math word problems with process- and outcome-based feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.608825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.608825Z digest=sha256:4a068e2633dd37e4f2bc05e80846d10f4a33fc67949e69372c0ed165f2535542

Observation 62d39061-9101-4d97-b714-4f020479f568 · outbound

This paper cites LiteSearch: Efficacious Tree Search for LLM.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension LiteSearch: Efficacious Tree Search for LLM

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.613397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.613397Z digest=sha256:a462e320eb0fb839cf99bf1aa5a2e7751bf875d3e6d84b678bf7f37059a78b9a

Observation 5be38ee5-e70f-4065-ad1c-08b6cc83a785 · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.618232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.618232Z digest=sha256:275e01ad88f28f0668fee7df730ad992a076d69c4f5c506e6274cf28e2caca06

Observation d746a5ed-51fb-4bb9-a5c3-bb0d9372ef2a · outbound

This paper cites Evaluation and Analysis of Hallucination in Large Vision-Language Models.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Evaluation and Analysis of Hallucination in Large Vision-Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.623635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.623635Z digest=sha256:ef2f92c03d4959510b98eef7f4ac51a612c746bd6636eed5fb11eefde944c270

Observation 730e666d-9e08-4508-8ac5-a044788a2e10 · outbound

This paper cites Amber: An llm-free multi- dimensional benchmark for mllms hallucination evaluation,.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Amber: An llm-free multi- dimensional benchmark for mllms hallucination evaluation,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.629411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.629411Z digest=sha256:ac5e5e8414d39458a36c9ca155224fa3027bb31365543b58c40c13453f690353

Observation 8f38ebf9-b39a-4cf5-adf4-6ba44d024104 · outbound

This paper cites Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites, 2023.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites, 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:02.933659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.634378Z digest=sha256:ccf68169a325aa3bf6d00ae891c8b004053f411a29836141d194dc3bedf0dca5

Observation ff318e63-3c8c-4698-a560-d3c7aeaee043 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:02.915119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.639100Z digest=sha256:a06f9b3dcee22df4726be96d7a73bf4596afb36d25a39d35f53e98c098b6478e

Observation aedc8502-7c1c-4ad4-95b7-891a54621db9 · outbound

This paper cites an unresolved cited work.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:18:02.895059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.644214Z digest=sha256:a9c62f54c14ed292d093433d9bebfefb56ab039f07fecdb373176ba85ea388fe

Observation 39a38205-86a0-41d5-8e18-fbdee6d4621a · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension CogVLM: Visual Expert for Pretrained Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.649362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.649362Z digest=sha256:ff3bcb77b3b831b21d991b46a880a88162d144a5d0b4b2b3f7afd7c87075d775

Observation 0c2dc636-e87e-4c1e-952e-e02447e94cff · outbound

This paper cites COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RL.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RL

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:18:02.017398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.657678Z digest=sha256:be132696e2378009a6c3d3b7dd40aabd181ae36298abf83307043cbd3d853a29

Observation 9771fe61-2df7-4e0a-b422-aa1e4f6af4bd · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.663857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.663857Z digest=sha256:6ccf8047fad28239061fcb02e07421907216ea0699740e6b6cdff660d96893ea

Observation f70a0c3b-0efe-4eb2-a558-a5bc29f97ab8 · outbound

This paper cites Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.670530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.670530Z digest=sha256:e0c17475fe47cbc1ab0f9a22956a4ca6e1bf576fd77c0ab0dd3fb0f070c898df

Observation ff63e70b-fe00-422a-bcd4-db3f26cf74c7 · outbound

This paper cites Mementos: A comprehensive benchmark for multimodal large language model reasoning over image se- quences.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mementos: A comprehensive benchmark for multimodal large language model reasoning over image se- quences

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:02.869239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.675588Z digest=sha256:b22ad3a649f86c36f99383f245f313df76c65bcb9dafeee91e56368ba83b1114

Observation e285b9f4-bcea-4119-a20d-f12ef37270f7 · outbound

This paper cites Simvlm: Simple visual language model pretraining with weak supervision.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Simvlm: Simple visual language model pretraining with weak supervision

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:02.844118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.679982Z digest=sha256:9a18200f469502063e3fad8da08c0ccaf9ab65b208dbcebe604dd3eea595301c

Observation c5a6e8a7-8e28-4820-a95c-b7590ce2bde8 · outbound

This paper cites LoTLIP: Improving Language-Image Pre-training for Long Text Understanding.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension LoTLIP: Improving Language-Image Pre-training for Long Text Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.684754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.684754Z digest=sha256:5ea1af4fefe94cb734ee18b2925abf795fe0daa3d06f82936428cbd7def4349f

Observation 1fdfb2ca-fddb-4747-b46a-c96efeb225c0 · outbound

This paper cites GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.689642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.689642Z digest=sha256:beb4b0f6bae887d30c5d3e86e636135f778a8581a940b457137cbca49c9585d7

Observation 0f563b2d-9a8f-4bb2-8887-07245d840361 · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Longvila: Scaling long-context visual language models for long videos, 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:02.822041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.694204Z digest=sha256:7dc290a16bb1387156533a988317bcf84c4dbab323a58045d42c60daec1ca88f

Observation 1aad679e-8437-4ac3-9025-5d4fd14683ac · outbound

This paper cites Qwen2.5- math technical report: Toward mathematical expert model via self-improvement, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Qwen2.5- math technical report: Toward mathematical expert model via self-improvement, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:02.783235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.698500Z digest=sha256:89fe114d1886c8aa65cbffab369d524f9e5d93269cf8522f85bf9e738406c665

Observation 8f567c19-0bf3-4de4-a897-560cf547d4fc · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.703890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.703890Z digest=sha256:9f714ec34b1a66a6b2127f471ff918deedeb4a80a0163ae37ed782005984238c

Observation e1d3a0bc-a400-4c88-8adb-7e38673599d7 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.708355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.708355Z digest=sha256:10f0dbffdc4d554211243b16d901a2ba5f5c98845c1c9b63156127a37debf9e8

Observation 106b08e4-2f6a-4260-a435-e521e6c9bc62 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:02.751846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.713282Z digest=sha256:6e75244dc33b5f60f59b7b304ae149ba89cfe841d996e0d7e801fe9b859ec0ef

Observation 4fc5ca45-ac36-4aae-8c54-a660053d2101 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Florence: A New Foundation Model for Computer Vision

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.717840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.717840Z digest=sha256:8e8c81198dd8d00469f7e19698a52bdbe29fc660dab88ae6715048974b84744d

Observation 408aa513-503c-46bb-ba62-0f3c7b5a7bf8 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi, 2024.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:02.731493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.723018Z digest=sha256:0ac53fafcf3f764a1df5c3fe8cdc6e42ba8c0b88595c2b889595c1e2c8b3e102

Observation 6b01f65d-f688-4e31-b814-775cf5a28ef1 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.727434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.727434Z digest=sha256:c1d98404a322222d5404fe5f64af656df2705a8292fda2be6518609a0e8fd098

Observation dd2b288a-b1a9-4041-845f-76fcceae8dc8 · outbound

This paper cites Beyond hallucinations: Enhanc- ing lvlms through hallucination-aware direct preference op- timization, 2023.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Beyond hallucinations: Enhanc- ing lvlms through hallucination-aware direct preference op- timization, 2023

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:18:02.710461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:18:01.731832Z digest=sha256:f9447f3c9febac71207c110ee3e3ae455d063d2c6431364bf093a66844a19126

Observation 13aee097-c0bf-44e9-93b6-0c71c731b307 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.735910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.735910Z digest=sha256:12ea33f9b9ce1b48b7d7277f036603f4a4a0a212dc75d32605a3f45bec0090e4

Observation 20b19f1c-6d04-4809-b93b-f02199b97e12 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Calibrated Self-Rewarding Vision Language Models

Reference 75

Resolution
malformed identifier
no resolver link, observed 2026-08-11T22:18:01.740237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.740237Z digest=sha256:75546fdfe85f9f4f87bee9cf8ddc12bb64edd2f46b12999da17561bfb50b2b68

Pith citing papers

Observation e1cacfed-d5a6-4c9a-857a-9f80a2ba00dc · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.089878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.089878Z digest=sha256:4a6b284d72dfb6461bafccc62072ba4cf14eb5793652073e314522b7b974aabc

Observation d5603a6b-9393-46dd-8c6e-5c5d0a5c7a09 · inbound

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling cites this paper.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.089127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.089127Z digest=sha256:976cd4d5b78013fa43d93e9016eab89baf7b3f55fccfaf296d66c1a0698510b4

Observation b69cd9cc-54b3-45bf-9283-819838cacd21 · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:06.014106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:06.014106Z digest=sha256:fe2d7daa7d84d56c62e15a097050ef4d43ea1ea9faebd5e26fb7e0018e68d79d

Observation 246c9b34-c7f0-4ca4-bee2-4043cdf08d90 · inbound

Mitigating Object Hallucination via Robust Local Perception Search cites this paper.

Mitigating Object Hallucination via Robust Local Perception Search Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:03.181743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:03.181743Z digest=sha256:9119938ba06b7e3adbb28f09b95f3a4c337694240b7c8ce3445ecdbc03d8930a

Observation b54b5691-eafa-49dc-87b5-550f54f6c437 · inbound

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding cites this paper.

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:18.661384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:18.661384Z digest=sha256:d40c124d3be123dd69668fd31f010996414a3dc258e165fbc4736dfd0b05f127

Observation d0da67b2-01d5-4d05-a158-b71c70f17dfa · inbound

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs cites this paper.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.741673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.741673Z digest=sha256:e526c076f4f69234dd7aaf34e797652464df3718fc82b37a5b52b73ab179ebfd

Observation 6add957d-08f8-48a0-9bdf-d3f310ccc30a · inbound

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning cites this paper.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.096140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.096140Z digest=sha256:0a19f6892982b1e01fbd3cbe01139721ec227612e85d94584dc6d08458c48031

Observation e74bdea1-5859-4b61-a8b2-633f4646b1b9 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:24:40.393917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T13:24:39.858355Z digest=sha256:91be6f2c44372bc73d63128943625fd20b38783806e89266904b2841f26eabf3