Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:18:01.740237Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 8 inbound Pith citation observations for arXiv:2412.03704.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:18:01.740237Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:18:41.089878Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T13:24:40.386203Z
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c8e51b86-0b81-4d14-805d-e682e575598f · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension https://openai.com/ index/learning-to-reason-with-llms/ , 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9a0e4367-0052-46e6-bb5f-fa2511edaac3 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e85f58e-88a0-4cb7-ab22-cfc79e54d146 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b621301-cb13-4419-a8c1-5bb2186a4602 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Improving image generation with better captions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2c120969-fdb5-4970-893e-4efbcaca9674 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06c1ee18-8307-4e8c-9f93-ada6a01830dc · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Transfer Q Star: Principled Decoding for LLM Alignment
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a4b06c-2c93-4cdf-a64e-07dcfb3bd8ea · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f8594860-4568-4fcd-98a5-7b1f76e26593 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4521f82-edf5-43e2-a32f-ca88a8b9f033 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Are we on the right way for evaluating large vision-language models?, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation afe682d7-db3e-4580-ab11-ef480f6352dc · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 681c3297-6802-4805-b162-e95bff58f360 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mitigating hallucination in visual language models with visual supervi- sion, 2023
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a353d46b-873a-4260-9dab-6cd45a1482e0 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ba321c9-89fa-448a-8b85-73b7439afbd0 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Training Verifiers to Solve Math Word Problems
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e41440e-4856-4ea2-ad43-5bba36692dc1 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling recti- fied flow transformers for high-resolution image synthesis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2db04c9-28c5-47ae-b547-e07b993b08e3 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1659d40-3742-42ec-97fb-5f673502d2e9 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fd60810a-6161-49f1-9ab1-3f5afd1aad28 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Temporal difference learning for model predictive control, 2022
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 12ddf2a6-6e41-4fd8-a2dc-360e4e7cca1a · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension V- star: Training verifiers for self-taught reasoners, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3f1690e4-1b81-4e25-9fa5-90509b98b4a6 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling up vision-language pre-training for image captioning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6e4539fe-dfb6-4e89-bea0-28536504ee53 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 916b628c-fa8a-4929-a6ce-f7054f106ace · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mantis: Interleaved multi-image instruction tuning, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 01a791ac-b095-4a8e-b28d-7419fd44d011 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69012f65-90d3-48f5-b595-e569857746ea · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Veclip: Improving clip training via visual-enriched captions
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e30b6298-7524-4f23-887e-acef171dc756 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e73729a0-9c58-406a-b897-4790b967c806 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Llava-onevision: Easy visual task transfer, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d36b022-486b-43cf-a045-d18933759d93 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Multimodal foundation models: From specialists to general-purpose as- sistants
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0d2c09b7-05d3-455e-89dc-c39153e4bf1d · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Common 7b language models already possess strong math capabilities,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation be12431c-c9e1-454b-89b5-b82695d36f8a · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models, 2024
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a14e9e-12e9-4a08-832c-9b279b9204ea · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Let's Verify Step by Step
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e4f717-0b73-43bb-b762-6220ff9e8d66 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Let’s verify step by step,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8814eb9d-d90f-4304-b8bb-3e3b10fdb5b0 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b18720-377f-4bbf-87cf-2c7631b8151e · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Visual instruction tuning, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19793b71-9131-4b85-bf01-21154a796e88 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Visual Instruction Tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e772ed4e-6e5a-4d0e-ba61-74d8bbab8b85 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mmbench: Is your multi-modal model an all-around player?, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23975e21-4eae-4fa0-898f-8c7b2c7e4fa8 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2121e488-736c-41bd-93d8-8694a3c22ac2 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Gpt-4v(ision) system card
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81191c14-d97b-4a90-857f-8d9bb695ccee · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Learning Transferable Visual Models From Natural Language Supervision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bbe461c-859e-45e6-800c-2fe59be2c1a0 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Object Hallucination in Image Captioning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16fb8192-cc40-498d-9fdf-55ca80f0de8c · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7308baae-38cc-4f4b-afe7-a3bbaceb799e · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mastering the game of go with deep neural networks and tree search
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 74df077d-7cc5-4228-b980-3674728ecef1 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29867786-e81b-48c7-839c-bc0b6d256b49 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling llm test-time compute optimally can be more effec- tive than scaling model parameters, 2024
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96572944-b38b-4991-9271-9f1a02ec53c0 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Aligning large multimodal models with factually augmented rlhf, 2023
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c167b5dc-8969-4f95-a088-0677866ad949 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Learning to predict by the methods of temporal differences
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0e8820ad-e1b2-49c7-893f-85e1f01cd99f · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Gemini: A family of highly capable multimodal models, 2023
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 26273c98-836d-4bfe-9810-fca71c780a5e · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Motion planning for au- tonomous driving: The state of the art and future per- spectives
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b40c4073-d9d6-41ce-b7dc-b47703ac6c4f · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afad285e-165d-49da-8831-af19fe27bde2 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Cambrian-1: A fully open, vision-centric exploration of multimodal llms,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22ab22c4-fa45-4098-96f7-b799195064fa · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Solving math word problems with process- and outcome-based feedback
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62d39061-9101-4d97-b714-4f020479f568 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension LiteSearch: Efficacious Tree Search for LLM
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be38ee5-e70f-4065-ad1c-08b6cc83a785 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension GIT: A Generative Image-to-text Transformer for Vision and Language
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d746a5ed-51fb-4bb9-a5c3-bb0d9372ef2a · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Evaluation and Analysis of Hallucination in Large Vision-Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730e666d-9e08-4508-8ac5-a044788a2e10 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Amber: An llm-free multi- dimensional benchmark for mllms hallucination evaluation,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f38ebf9-b39a-4cf5-adf4-6ba44d024104 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites, 2023
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ff318e63-3c8c-4698-a560-d3c7aeaee043 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aedc8502-7c1c-4ad4-95b7-891a54621db9 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 39a38205-86a0-41d5-8e18-fbdee6d4621a · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension CogVLM: Visual Expert for Pretrained Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2dc636-e87e-4c1e-952e-e02447e94cff · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RL
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9771fe61-2df7-4e0a-b422-aa1e4f6af4bd · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f70a0c3b-0efe-4eb2-a558-a5bc29f97ab8 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff63e70b-fe00-422a-bcd4-db3f26cf74c7 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mementos: A comprehensive benchmark for multimodal large language model reasoning over image se- quences
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e285b9f4-bcea-4119-a20d-f12ef37270f7 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Simvlm: Simple visual language model pretraining with weak supervision
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c5a6e8a7-8e28-4820-a95c-b7590ce2bde8 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension LoTLIP: Improving Language-Image Pre-training for Long Text Understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdfb2ca-fddb-4747-b46a-c96efeb225c0 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f563b2d-9a8f-4bb2-8887-07245d840361 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Longvila: Scaling long-context visual language models for long videos, 2024
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1aad679e-8437-4ac3-9025-5d4fd14683ac · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Qwen2.5- math technical report: Toward mathematical expert model via self-improvement, 2024
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8f567c19-0bf3-4de4-a897-560cf547d4fc · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1d3a0bc-a400-4c88-8adb-7e38673599d7 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106b08e4-2f6a-4260-a435-e521e6c9bc62 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4fc5ca45-ac36-4aae-8c54-a660053d2101 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Florence: A New Foundation Model for Computer Vision
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 408aa513-503c-46bb-ba62-0f3c7b5a7bf8 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi, 2024
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6b01f65d-f688-4e31-b814-775cf5a28ef1 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd2b288a-b1a9-4041-845f-76fcceae8dc8 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Beyond hallucinations: Enhanc- ing lvlms through hallucination-aware direct preference op- timization, 2023
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 13aee097-c0bf-44e9-93b6-0c71c731b307 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b19f1c-6d04-4809-b93b-f02199b97e12 · outbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Calibrated Self-Rewarding Vision Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1cacfed-d5a6-4c9a-857a-9f80a2ba00dc · inbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5603a6b-9393-46dd-8c6e-5c5d0a5c7a09 · inbound
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69cd9cc-54b3-45bf-9283-819838cacd21 · inbound
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246c9b34-c7f0-4ca4-bee2-4043cdf08d90 · inbound
Mitigating Object Hallucination via Robust Local Perception Search Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b54b5691-eafa-49dc-87b5-550f54f6c437 · inbound
What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0da67b2-01d5-4d05-a158-b71c70f17dfa · inbound
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6add957d-08f8-48a0-9bdf-d3f310ccc30a · inbound
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74bdea1-5859-4b61-a8b2-633f4646b1b9 · inbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.