Pith. sign in

Paper Citation Record · LEDGER

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

As of 21 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 7 inbound Pith citation observations for arXiv:2411.18203.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18203 v5

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:27:33.494574Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:20:20.174537Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:36:16.802423Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved55
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 034a4a1f-25b6-421c-ac16-9da194944231 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.082475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.082475Z digest=sha256:e626426908e48b595e1ff5acf1cb0e9da4f03cca45c23b5bc2b314e2abc9f15a

Observation 8d605576-d546-41e7-a892-58bb37b3040a · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.088452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.088452Z digest=sha256:eff114eb997b250d894f04dd9dfe53266568d9319718d6a12df132e27efc5ce7

Observation 1f148c67-6662-4cdc-bf32-6d7fac83682a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.094064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.094064Z digest=sha256:466d4b3d5c26944cc48affe68c735f5c84d84beb3c1da9f8403934ca8f7d7899

Observation 2167781a-0df1-4004-8d67-5b91be868b77 · outbound

This paper cites A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.101376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.101376Z digest=sha256:96805b3eb450bfac5458b1d2f2fd5d16fb125481df2863964959a547d78be179

Observation d8221068-bef4-43bd-b7a2-42b5c0ed0d0f · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.106768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.106768Z digest=sha256:0f734bb92bff0c46ce5023ba6516e698777a2ef7d15dc4e07cb327c5b56c46f7

Observation 9d0e1993-dd2c-499c-a8e4-8e3114a8f8e8 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.112655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.112655Z digest=sha256:d0346a407714fd3555336551112eab584490ac3d05e8291157c85620d7ca848e

Observation be18d887-8108-4785-aca9-1caeec776a58 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.843995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.118009Z digest=sha256:f6ca13f50fda6aca020fe033f342f797c185a472042c6203ec56591f4da6dd10

Observation 92ea7bac-5f81-481e-a256-205d9dfafc9d · outbound

This paper cites CAST: Cross-modal Alignment Similarity Test for Vision Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning CAST: Cross-modal Alignment Similarity Test for Vision Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:27:34.137559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.122826Z digest=sha256:8dae3aae59ea310f901ae420a8055db4bee0864be7e47f06b7ea7f11f394442b

Observation 979b5fa9-23a3-4333-ae07-42be3cf973d2 · outbound

This paper cites Gemini-1.5-pro, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gemini-1.5-pro, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.826294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.128478Z digest=sha256:594c2ca969492a78984094d3362afb6d31d3c123126142d7de4cdbd3a23ddb39

Observation 95a58b8a-9adb-4645-9eb8-49aa8962c5b5 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.133335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.133335Z digest=sha256:e3e7acc5d1d3dae5b33f6399147dda4812ad654b3e73a1197bd0f5621b4cafe7

Observation 0b954e3b-be66-4f36-a2ed-e11590f3c993 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.138054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.138054Z digest=sha256:c78d110e6ed342744e42cd6baa896acaba545c2218afae6de8bde2e52560ee31

Observation 6be69ab7-8e1b-4659-ac89-bf5ecb436ff1 · outbound

This paper cites Chatglm: A family of large language mod- els from glm-130b to glm-4 all tools, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Chatglm: A family of large language mod- els from glm-130b to glm-4 all tools, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.809098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.143614Z digest=sha256:93464126873b6d57e0c1e29b6b1a6a1bb0316cd2427ace5355f15169c50c5fa9

Observation c9079af9-cc7f-4116-917a-db2bbdc4aea0 · outbound

This paper cites ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.148992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.148992Z digest=sha256:2e643602a8f237c8f1749e368f8ffdf04eb06a7e5627a00ad16e438b3e7a2f99

Observation 3731d655-9dc0-467b-a652-6bc6913c5e71 · outbound

This paper cites Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.153845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.153845Z digest=sha256:2a4cabc0c36f29e8544f1773d6267b3247c57ecd7f190ca52ea77dad752b5dad

Observation 5a8f6af5-373d-4af3-b4b6-b2e5fd8a749e · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.158423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.158423Z digest=sha256:7b06319e4167d8f436a22ad01a47c8ae8b246c78105760dd1b103d2c489c9515

Observation 6860837d-de31-4a5d-947d-ed08abc4ad61 · outbound

This paper cites St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.163296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.163296Z digest=sha256:badb15ef8eb4891aa699b7a071296a3732683b2b8a76c8b210463e32ecef1214

Observation 954f6cb4-c8c7-4c1f-9733-fc183994a71e · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.167657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.167657Z digest=sha256:23aec4ebb7a98d8ab17f1591045b714a0bfd01beb45684d23cf39c6e71c6a301

Observation d5c2c2d9-f8ab-423f-b1b2-7c8658dd7850 · outbound

This paper cites GPT-4o System Card.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.173463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.173463Z digest=sha256:6b93a018b8f5c1d858004b2e8e325e09e5b3bfa5a61aab61b38d8b6be39a7b3c

Observation efff575c-602d-41a1-8d0a-52067462f1e5 · outbound

This paper cites Vad: Vectorized scene representa- tion for efficient autonomous driving.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Vad: Vectorized scene representa- tion for efficient autonomous driving

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.178940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.178940Z digest=sha256:2c24c8eed1acaaf1d0b017e691b15cbb47cbb22390a2e4a40b8a3c8efcdff7fd

Observation 31478b72-4aa0-4807-8de0-ab17a6049e56 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning VIMA: General Robot Manipulation with Multimodal Prompts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.183592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.183592Z digest=sha256:01f48acf8a4ad61b77c1b6d3f8e0210c894922379b67836f54c22bf4b7552e43

Observation 8b38f95b-ee23-47cf-ba57-cae2dcaf62b8 · outbound

This paper cites Large language models are zero-shot reasoners.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Large language models are zero-shot reasoners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.756313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.189253Z digest=sha256:dd62ca1faf226787c43206b62781cd1407730eb0256e8c3e20ffa8cae8617814

Observation 6bb5eb8b-6966-45ff-9ac3-634d3d9238f2 · outbound

This paper cites In-context Reinforcement Learning with Algorithm Distillation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning In-context Reinforcement Learning with Algorithm Distillation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.194094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.194094Z digest=sha256:b024e4bf710165850bbbfaa05e4ab6be8780241beabc69a9a4e633c19874dce4

Observation 1e646ab3-d26a-4239-9865-26d51064c042 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.199026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.199026Z digest=sha256:7e8e042de92e618a1a9211750cd7f9e8b5dab5c12f2f32c315ca9c833b072153

Observation a01a592e-44bc-406f-a3c8-a5ab89e09ed0 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Silkie: Preference Distillation for Large Visual Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.204022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.204022Z digest=sha256:a64c0c07b59842d591addaee95f0f2381823386279644cc59f57ad45bcd0afd0

Observation 9f764ecb-cd88-4e22-9e05-d9beeb962bcb · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Evaluating Object Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.208990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.208990Z digest=sha256:2a0e4e1c3235ceb5aba4d02a5fc77bae2feb43984cd896325af2a39cc36e6bb3

Observation 52f9c951-d830-489f-9ce0-fb42c12adb18 · outbound

This paper cites Let's Verify Step by Step.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Let's Verify Step by Step

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.214333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.214333Z digest=sha256:7d4dadee4ceb5c041335238b8ed58118103e74c994706fae7da318fc4e0a5bd9

Observation 048b293e-b7d8-4666-bcab-cd50cb634d64 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.740575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.219132Z digest=sha256:13ef5c5361712894e55a3f1f2c42ad841f39d6fddc9a4b7d55435095799dccda

Observation 28aa8fb3-447d-4c19-9214-0325a9f0d193 · outbound

This paper cites Improved baselines with visual instruction tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Improved baselines with visual instruction tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.223939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.223939Z digest=sha256:d0fbb97901b0f71bef16841f89aa1a307455de39d1630e0ca806093c96fa1a49

Observation 8bf43c82-feae-46e3-949b-2f7e1a3902c6 · outbound

This paper cites Visual instruction tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Visual instruction tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.228945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.228945Z digest=sha256:aee65da6ec4a24ea7eb50dd789cffe7030013268d4328a5c9cf834f30dc3f0de

Observation 422cb8f1-5374-49ab-8131-14a6d80f8491 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.702876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.233647Z digest=sha256:02f341f45183693b81d05aa0638a143ca9998a63b7408d96257dfa48d4bb4860

Observation a0b62837-e951-479b-bf63-bf0766baa20d · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.238871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.238871Z digest=sha256:b27cedc18d4c4247d734862d216369adc4f920cc82ddd11659d9c1d46f08bc03

Observation 06df21f4-afa2-4141-b501-5b7627c19648 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.685242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.244301Z digest=sha256:b51d612c2f7b3a3aad4825ffe49929057a90860d6a2664dad38ff4e79af387d7

Observation 87ede84d-8f10-41bb-a46d-ac44467b0c45 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.248747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.248747Z digest=sha256:2ad6dcb1e676eac4d04e9f0e0e08997783bf5926ce29c59b954236925304b22c

Observation 6495c531-cf02-4d26-9b78-8d16ed8d249f · outbound

This paper cites Self-refine: It- erative refinement with self-feedback.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Self-refine: It- erative refinement with self-feedback

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.652549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.255007Z digest=sha256:ee03b11ec6cf21c806d7db90d56ff995ae25551914df0c22d74319174a39715e

Observation 8d63ecfa-947c-4889-b4fe-80eec2bde86e · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning LLM Critics Help Catch LLM Bugs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.259690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.259690Z digest=sha256:fd0c336352bffa208e8b129021863c0b3d2bcf7f2097d5207279a1640118ef58

Observation c2bc84c7-b20a-4dd5-bcf5-f641e5a9f4e6 · outbound

This paper cites Llama-3.2-11b-vision, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Llama-3.2-11b-vision, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.633589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.264863Z digest=sha256:a42b68dcc2b641e0269b4a8c0b39628459c506212072b33c5b29bdb759971751

Observation e3d8be13-5f55-4c14-9033-84d6af8a8726 · outbound

This paper cites Rule Based Rewards for Language Model Safety.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Rule Based Rewards for Language Model Safety

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.269702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.269702Z digest=sha256:f1833f6e46679b1ba42b6aee8653cf1e92d20d15c19564b3ec004616b015a538

Observation d1ecdada-f367-49fc-b211-fecd9b6778f8 · outbound

This paper cites Gpt-4v(ision) system card, 2023.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gpt-4v(ision) system card, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.615390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.274830Z digest=sha256:0e760d4be8d53a6e6b3ecc1805148af44732e1efea1946dbf0664984f27edc06

Observation ba6bca92-4df6-40ab-8c03-ec2b0de626d0 · outbound

This paper cites Hello GPT-4o.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Hello GPT-4o

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.597694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.280224Z digest=sha256:0b9ad38b2ba62acd7e22a92f3a5096ae2821ae2e3630e799b32c0a714decda75

Observation 4ab1b5ae-9ce7-4de7-bf41-ed108ee4c08f · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence,.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gpt-4o mini: advancing cost-efficient intelligence,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.285300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.285300Z digest=sha256:f8f78e21a843204137aa16c8edcb9d603e92952b9fe094642c6f8a3c4838a00b

Observation 58cc6361-bb9f-4fd6-a1d4-e57357a9388f · outbound

This paper cites REFINER: Reasoning Feedback on Intermediate Representations.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning REFINER: Reasoning Feedback on Intermediate Representations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.295813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.295813Z digest=sha256:0675692f426bdf2f76be21708bbd1f36ce33fb8b696a594c0e12c8b78da56d48

Observation d29ed5ec-50d7-43f7-9572-3f24f6753b9d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Direct preference optimization: Your language model is secretly a reward model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.552174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.301581Z digest=sha256:5738cb6bac4adad398dcc9e85306f205bb6f11217ff5ecfbe1d4e08f9c0193a1

Observation 78046100-131f-421f-9de1-8f0bbfb6e49c · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.533861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.306734Z digest=sha256:323d78889986bff4860aea1cd799581115749ec5559aeefbf848068c58bb053b

Observation e0cf320e-7c19-4f4f-bc7b-83ce7c498fbf · outbound

This paper cites Learning to summarize with human feed- back.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Learning to summarize with human feed- back

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.517214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.311992Z digest=sha256:79ce7e6d4ab8665e98548629e9f05321a1e02eb4a63ae53dde48b3e685f4dced

Observation 3c94c04f-de9e-4a45-b91e-7bfca6409f05 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.316975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.316975Z digest=sha256:dbba4058e6927ec389d79a1876e69582139fc1ed7700d4ce324771d791d39889

Observation 266575a5-ea79-44cc-91ad-8395ae060719 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Policy gradient methods for reinforcement learning with function approximation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.498472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.321810Z digest=sha256:023d896ba76c0997faa04a177ca67d97775d5e807d8c8d56e38fc74dabd3d585

Observation f02a630f-3b4c-44b0-83de-099917c43c33 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.326629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.326629Z digest=sha256:e5cfc1ff810ab91c876bf859e53f49fa7ee51f2717e099f86b27e00400fe6bc2

Observation 2bbdabe7-3460-4516-b7d9-ab3ac4b6feea · outbound

This paper cites LLMs cannot find reasoning errors, but can correct them given the error location.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning LLMs cannot find reasoning errors, but can correct them given the error location

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.331621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.331621Z digest=sha256:e7d98cabd9756a8ebbcc84dd1a0972943efaab2b94cd9bf4fa13a5671905b23c

Observation db7a87f6-8cf6-4d55-818b-9e8a057a47ef · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.337313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.337313Z digest=sha256:eda2d759f8b5f53f8833e4da80602a82bfaafc62fbd6f56544cd323582f917f2

Observation 17cfad8e-0140-4bac-9f76-730c358064b0 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.342823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.342823Z digest=sha256:93a4e40ef549c13ff4e25505ae808d39daa498174ed30265fb23495db1166498

Observation 5b9e669a-383e-4016-9fcc-099afb58f0ed · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.348467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.348467Z digest=sha256:6acd6d8a8109d97a3091fb2cb4750c069279fcae6186573b2d8b4566b0c5cd13

Observation 917ed366-babe-4e96-b341-d5776162cc89 · outbound

This paper cites Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.353133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.353133Z digest=sha256:52abef4a868f090b9e8d31fb8e39c88ae39a306fcb3916f92a04c81d026214a2

Observation bfaeb753-cdcd-4070-aee9-6197022b293e · outbound

This paper cites Grok-1.5 vision preview, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Grok-1.5 vision preview, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.464735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.358203Z digest=sha256:518c8ced6d97ccd873685a28d074be746b72fffdb3f54a2b6f52e6d8bb4f7bed

Observation 710f667f-063a-44ce-8c8e-4465ef4a39c5 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.363118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.363118Z digest=sha256:c1492e3ec89d79afdae4773f4299a669bc7a74eb0d5450ad143e246dcca2b333

Observation 1087fa5c-e965-4048-83aa-9e4506dd8fe0 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Tree of thoughts: Deliberate problem solving with large language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.446738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.368550Z digest=sha256:2301bf2bdfa2ccbf6adc95a36771dda69c3d7809676b75142d45693a4ff604cb

Observation 0625d111-4cbe-4aba-8ce7-1f3acee4675c · outbound

This paper cites Learning From Correctness Without Prompting Makes LLM Efficient Reasoner.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Learning From Correctness Without Prompting Makes LLM Efficient Reasoner

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.373868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.373868Z digest=sha256:d9716ef7e2833b423249043a5d64bf736cf72a1b1d7906d728dcd1e5006467d1

Observation 10c014a5-3ca3-4a3a-b5d6-c94b48c9adc5 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.379680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.379680Z digest=sha256:dbbda0875533f235e184a4b87ac1717b6418804e9b777d008db2184a27345556

Observation aa5b0aad-32cf-48c4-b16c-0316c66d5e4d · outbound

This paper cites Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.428259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.385690Z digest=sha256:f0f89659ec47b3aaeeb728e1ea25b52638dd7d00160314406c229f5c6e6ebdca

Observation b9a1de13-d079-48e7-ad42-73c40f21d99b · outbound

This paper cites TextGrad: Automatic "Differentiation" via Text.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning TextGrad: Automatic "Differentiation" via Text

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.390355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.390355Z digest=sha256:2c07c6bc1f7dc95509b6f3b419db1cd0f84fe40ce83bbe3eae1e528dea478a64

Observation bde391a1-1fb1-4797-91be-ec23c524300c · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.397485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.397485Z digest=sha256:f849fd4a3a905b4ff70ad8fa280d05b19299e48abd6c34ddf2108eca6c3729ef

Observation 4c064ee6-a811-4e7e-87f3-006990c4a0b1 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.403459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.403459Z digest=sha256:6c2a845fd56d3131180b6befed50434670ccf3e9fca9f005c484d731e6c89909

Observation a8d72bef-65e9-4fce-9748-e725d2b3ee7a · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.408895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.408895Z digest=sha256:f67679d55d24b3ed955597d8c54b40da85854ab7ae350cd91118ced35c63aeaf

Observation b0448ac7-50ff-4f4e-92e1-a6e1e4d13b76 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.414817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.414817Z digest=sha256:a61b3046ba2590ac2559ef585ca17e81254a3949c776c6472fb80fc84e7ee218

Observation 191d3506-af80-45c3-a24a-ad877de06601 · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.419523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.419523Z digest=sha256:d876ff2db7a5e4eb243de229277a9d974e396b1b4302637c36845d457e6c1b5d

Observation 763db4c7-4943-45db-9dd3-d365522f91a6 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Automatic Chain of Thought Prompting in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.424887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.424887Z digest=sha256:9337cd052762eac32d806cb24e22b0ca821e5ada45fec459d8c0aaa68806064b

Observation 2cec8827-945a-465a-b57d-ed01b74ed74b · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.430241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.430241Z digest=sha256:714dc6c98cefb491e788204ac10da964275fae4a70a88cdc3f73e0ee6285bc93

Observation 14067b64-e3d4-4320-8f78-261170d301f2 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Calibrated Self-Rewarding Vision Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.435515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.435515Z digest=sha256:aae2e785c7870fa7f46527ca76d0c71a4fee1d9eb7f624d7bbfc99cb02f12f64

Observation 221d1c84-03bf-489e-a144-4c5434667161 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.441121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.441121Z digest=sha256:0e94b93b0ea554f1afbcb94f5b4611e4b96a476b748e335f225e4336b527de43

Observation 6a56b08c-c0b5-4c9d-9001-1f9b8121916e · outbound

This paper cites VGA: Vision GUI assis- tant - minimizing hallucinations through image-centric fine- tuning.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning VGA: Vision GUI assis- tant - minimizing hallucinations through image-centric fine- tuning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.410929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.446544Z digest=sha256:de97976effcf3bb4afc748adcef1e8b6bc3e4ca96652d4be6042a20a7203a508

Observation e3b08c04-0c7e-447f-baba-6f68ca9d0e44 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.390545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.451714Z digest=sha256:5ebeaf87aa8434867a5216f48f9c09d002f009eb780f69e0434d26f0324d4aa7

Observation 99993b3a-981e-4019-bf0f-125bb30f6618 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.374596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.458625Z digest=sha256:95564e3162a193d6d6a4e0b3e07602a5f88f86e9828372650068a175d29eae30

Observation f0b1f474-4668-44e5-b06c-db83442fcaa3 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.356708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.465299Z digest=sha256:7b4af4c73f064cade8387c31391d715e225875aac2a9a137ce5b2081038312c0

Observation 48cc2e8b-a86b-46e1-a840-d6ab3448261e · outbound

This paper cites For preference-aligned fine-tuning, we utilize Direct Preference Optimization (DPO) on 29,012 samples from the critique- VQA dataset.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning For preference-aligned fine-tuning, we utilize Direct Preference Optimization (DPO) on 29,012 samples from the critique- VQA dataset

Reference 74

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T11:27:34.332142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.471122Z digest=sha256:4a87a727357cffbaeaaf48204c395466cc42a3d7519882a50947193eae38937e

Observation 19a48ded-338b-493e-a1e1-4d160a218afe · outbound

This paper cites In this section, we will list out the hyperparameters we choose for evaluation.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning In this section, we will list out the hyperparameters we choose for evaluation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.311708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.475936Z digest=sha256:30669d56084aca199cde5eaa21d20163b94e5d9929fee5cc4280991017902dff

Observation 470ff0d4-673e-42dc-94da-807a7963785c · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.295105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.480557Z digest=sha256:a5c5e24f2b3231e5d561fb87c742e9ca1cddb2bc0085aba8b1b35ea0b64485ce

Observation 1e61a02c-5122-4f37-a31c-9e03ba72aab7 · outbound

This paper cites You can find them in Figure 6, Figure 7 and Figure 8.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning You can find them in Figure 6, Figure 7 and Figure 8

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:27:34.274023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.484764Z digest=sha256:807752b0d284f6efdfaff7fde2b9c196adfbcdbb3e17b469cc9c4d71b4711bc8

Observation 1a395061-503d-4360-aaf8-4d1ae46dc235 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.253077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.489216Z digest=sha256:aaf295903f177d4268a1c105d985b47685b8a9c2341dcc0f626ad650f381729c

Observation e39a1e81-85d4-441c-b558-f7f657f6c09f · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:27:34.236407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.494574Z digest=sha256:92c873ce368df002f55ddab11704e47202edd2a4fe9fd778f262e982d9156529

Observation 84196f46-7d29-4e07-98c6-98ff17b6ef59 · outbound

This paper cites an unresolved cited work.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T11:27:34.569322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T11:27:33.290232Z digest=sha256:7b66b30fa2e3ccfa3a4cb87a51c00c0188f28d097f59a754cdd887cc1fec83f7

Pith citing papers

Observation f7ccdbab-a1c1-489f-8523-1d22700a6232 · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.308563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.308563Z digest=sha256:fa8b2b9494bcd4a8cb06c0c8fc1ca6422aecd69e89feaa1cc8ffc5ba3215b527

Observation a9c32915-a540-443b-8972-dd0a4f20f100 · inbound

ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification cites this paper.

ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:20:20.174537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:20:20.174537Z digest=sha256:afd52067cff24828c101f55f5dca2d4e6c8e540d5f0d8e2580d9de1eebac93b3

Observation 3dd5afbd-cc1f-4a06-883e-04f55644e95a · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:05.875185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:96e8ae07bdbcb024f12916bc363da683d08736c105ff0ec95fbb72e62e745daa

Observation f0639fcb-0b98-48d9-b284-46d6eb8d4c83 · inbound

Test-Time Hinting for Black-Box Vision-Language Models cites this paper.

Test-Time Hinting for Black-Box Vision-Language Models Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.061496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T21:16:45.748116Z digest=sha256:ed3253d7e3c37426eae4cd6d497354ed603bbbec64001061b7913a7f3407d4fb

Observation c33ebc55-f171-4fb7-a660-b1458662dde5 · inbound

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models cites this paper.

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.803674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T15:19:26.983760Z digest=sha256:67b65a1aca8c979817162266281f94115d98c461c1f16ab6fe6f957e903f3754

Observation 534359d2-0f71-4412-8ef1-5ec5e2c6d94e · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:16.664589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:16.664589Z digest=sha256:2c79c5a0278b5c0a837aa86215d1d8f36fdb2f0e17b9c28b0822c4cabf532972

Observation afc8d2b1-b606-46cd-840f-99b58920453f · inbound

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus cites this paper.

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:56:01.975427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:56:01.975427Z digest=sha256:465040ab6dc732c86cb950d2480c48bead8a53a8860561b52db7d6ea8742c0cd