Pith. sign in

Paper Citation Record · LEDGER

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training

As of 8 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2506.13888.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13888 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:34:03.367306Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T04:33:36.634359Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T21:41:16.983561Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71f73914-ffcc-4ad3-bd6f-45c80ff82150 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:53.663253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:53.663253Z digest=sha256:98ad80efbdb62817f929ca1905c275e973e1caa90f2f5bbae840db98957f5285

Observation 5079a9e2-c7f8-4ffb-9c4a-243bfe007e95 · outbound

This paper cites OpenAI o1 System Card.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training OpenAI o1 System Card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:53.903503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:53.903503Z digest=sha256:2cdd1ea8f378de8f47439f75d6ae4b5b8f5f0cd84bfa494d81f62a889ff57072

Observation cf97bcee-cb13-452b-a8e5-636f7213a605 · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:54.034506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:54.034506Z digest=sha256:94f635ae9f0239ec111445746bd5b49446248524d44c7a8c48a8f7b7e702106d

Observation 81c9655d-a4f7-480e-9544-4e8fff1455be · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Aligning large multimodal models with factually augmented rlhf

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:11.216046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:54.154868Z digest=sha256:c2890947c13fee0daa91791e8476209c673f5134da05a8a08f7073abdc6c80d2

Observation 5e70e07a-1c5a-468c-b233-6d55cccb9a60 · outbound

This paper cites Silkie: Preference distillation for large visual language models, 2023.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Silkie: Preference distillation for large visual language models, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:10.902201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:54.359009Z digest=sha256:f190db62c95fa5784efb49f70148df515c58e0e8b047e5a290c4a2ebc8851c95

Observation 92dc206b-2f27-4936-aa35-eeeed62849e3 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:54.489101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:54.489101Z digest=sha256:6c41f72790087b522062182a040ddc48871718c6cdff7ccfaf3f84454c9c13d0

Observation 6e16397a-fb4e-43db-90b2-7bc3c5970f64 · outbound

This paper cites Rank analysis of incomplete block designs: I.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Rank analysis of incomplete block designs: I

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:54.584748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:54.584748Z digest=sha256:19e16504f22ef39097d1eed5aede2a331a4c25d4be42d297ddd727e09747bfae

Observation 19d226a1-faaa-4712-8c19-de7c679ac146 · outbound

This paper cites Strengthening multimodal large language model with bootstrapped preference optimization, 2024 a.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Strengthening multimodal large language model with bootstrapped preference optimization, 2024 a

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:10.724456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:54.714834Z digest=sha256:1fb1d949bd57a25f9b384ab24c39aa4a8b1452f75894ee23bf60d7f14ccaf9ee

Observation 92bb1aed-1ce9-4191-ac09-1bfdbc1d0dee · outbound

This paper cites RLAIF-V : Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training RLAIF-V : Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:54.836665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:54.836665Z digest=sha256:52c3edb60ae1381f788be1c368d1ce6b286dd0d2dc80205f8d45009274d2829b

Observation 8407a231-d41b-4fa5-9edd-13a16d08b29b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Direct preference optimization: Your language model is secretly a reward model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:10.525501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.007073Z digest=sha256:0bab8e00e6222079cb42cb8120272ce2374403011b04f132f1a48572c998c73e

Observation 85f44ca7-2d2b-438f-8211-0fd8af44114f · outbound

This paper cites Self-rewarding language models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Self-rewarding language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:10.219552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.131233Z digest=sha256:463193a453c043acff21f8aadd79deb356f76d13b2330b7950b937492e76328e

Observation 897f665c-b0f7-4224-82fe-a734ff476882 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:09.840092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.313774Z digest=sha256:a395fa1cb5240585c0330da6e945333f6026811337a7599344288f9cf7d6811e

Observation 43432bc8-86d9-40e3-8f00-ea415d2bb00c · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:55.414768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:55.414768Z digest=sha256:1c50171bdf055a88b10341f77503041932e16f0d5211e66327a1e1e2ccc04ab2

Observation aec22b80-3c18-4d55-b9d6-c941a1c09b7c · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Star: Bootstrapping reasoning with reasoning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:09.535629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.542063Z digest=sha256:ad1bb87bf70e4ce4a934c54cb4a6357223e6f89c450c40ef80d14ea4b8caf537

Observation ea707d35-2e93-4723-aeed-3f9c65a8d884 · outbound

This paper cites RAFT : Reward ranked finetuning for generative foundation model alignment.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training RAFT : Reward ranked finetuning for generative foundation model alignment

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:09.221417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.710576Z digest=sha256:d5c5760631c654692beded5abd4675275e3b11477392f9ecba52c7d6006dee75

Observation 9d3a3448-bd1f-4dd3-9b1f-e59ec1e1b857 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Rlhf workflow: From reward modeling to online rlhf

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:08.876205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:55.944239Z digest=sha256:a2fcf3b2cf54aaf407bd008a442199c67f0de425845878c56c41e3b0e8d5384e

Observation 335a85c3-ee6c-4059-8461-65bf81add525 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.124755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.124755Z digest=sha256:3b33ee6bcda64e82c1a5bf0a9b01f8f1d5be17d426824cc1f836d1b64d5ef3df

Observation 8ff8a13b-870f-4430-9788-4a5a244680df · outbound

This paper cites Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.251004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.251004Z digest=sha256:6970ecb3232000b67eb0d9d7bf4cdb9a215cd39049bb6c50af27b8e32ed03976

Observation 7d51a540-97a0-41d0-9202-02ef0d0b3ea4 · outbound

This paper cites AIDE: Agentically Improve Visual Language Model with Domain Experts.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training AIDE: Agentically Improve Visual Language Model with Domain Experts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:34:04.544547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:56.525727Z digest=sha256:d26575b3280ea01c56fe602390c0e5bac3bb4c863e8accdbc23e158333a00fc8

Observation 5f4fd0b5-7eee-4cc1-92ea-93643c245a82 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.665520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.665520Z digest=sha256:723070c3cf9669ace2db4c9234423422195c6d7a69b709e5913c510c26cdf699

Observation 4dc8e471-1f47-4447-9f08-d0e01cf6afef · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.770706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.770706Z digest=sha256:77207ee743eab13fe12a4a3d78264724ec682b4860328093fadad811806cfd45

Observation 17375a3c-39b5-4c5f-93b5-1113f5dbca4a · outbound

This paper cites Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.851632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.851632Z digest=sha256:7fb5816058e44007acc03c209830ae4e21717abdb582daf1789cf526ce664695

Observation 712e93bf-a4be-4b77-8f3b-32299be42706 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.965198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.965198Z digest=sha256:1b2a3c5ad0699f22665ceadfb2c674393e572f0e3b61df47af44ada03c153bc0

Observation 97ee5a45-5a24-47aa-8ddd-27be7773c5e7 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.076467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.076467Z digest=sha256:ea57db4f66e2275c8d46dba25e4922bfdf70ee1d8288aae630e74f6a64547dc2

Observation 3af1a78a-4148-4762-8f6b-41e1dcaeab74 · outbound

This paper cites Chi, Quoc V.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Chi, Quoc V

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:08.485421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:57.145526Z digest=sha256:39abeb23b27e5c4da7b2695ed13310a9ae8f879a2cbce18076b922271111e252

Observation 7300f65a-3223-421e-b689-8e8ecc2a5f18 · outbound

This paper cites Iterative reasoning preference optimization.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Iterative reasoning preference optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.337884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.337884Z digest=sha256:2a145afaf62eb156bc12b25f89c9a046839bb3de4146725bb87fb8c7fe47c5d3

Observation c92ff6c1-a64a-4505-b622-bf0e88232d6a · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2023 a.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2023 a

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:08.231425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:57.515024Z digest=sha256:79f2d9a17dd6f00eb9ca0868c0c78382941b11539fb05d01217da5fc3ebd251f

Observation 2fda5150-1f12-4cfd-a2f9-b595027bccb9 · outbound

This paper cites Detectron2.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Detectron2

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.605814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.605814Z digest=sha256:c3f08ac3c6468305a422a3fc6a9c6d13a11c10f71e9aa0369670a603a45b4c66

Observation 057fb7c3-9c81-49c3-aff1-b0695aace441 · outbound

This paper cites Proximal Policy Optimization Algorithms.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Proximal Policy Optimization Algorithms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.696410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.696410Z digest=sha256:96e63fc395561fc9ad3b50a1882709e32eac3053b64ae85f93f517fc54dd66ea

Observation 9234ae8b-c117-484c-82ca-a7c707cd8634 · outbound

This paper cites Training language models to follow instructions with human feedback.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Training language models to follow instructions with human feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.824861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.824861Z digest=sha256:0584e0ece385775ee24f90fe60ce92d5ce6c7fd1ce56d0f3d6c6df402ee69ca1

Observation aeb1c785-afd8-4663-972e-468a808ab679 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:57.925619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:57.925619Z digest=sha256:6ab5d80c0478d6e6453fb4273c89d85dbfd5b899847e55e92189da04162800c9

Observation 16ab5fe5-b16c-4bf7-adaf-8a74df19abd2 · outbound

This paper cites Hello gpt-4o, 2024.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Hello gpt-4o, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.089473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.089473Z digest=sha256:9ab673355206f4291bcc8ec1f9c89e32017b70bcf86c51356b2d6d8c735fba07

Observation 3aea31ff-d467-4ea0-a473-3a36fc822df8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.204744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.204744Z digest=sha256:f7565a61211f1ac89c1f886cfecb9497d2848a4f484b115c301b3658a1e7f47b

Observation 4f756b18-5c6c-41e2-a79a-6dc72c1eeef7 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2023.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Gemini: A family of highly capable multimodal models, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:07.881904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:58.323913Z digest=sha256:9b4fc3bb7ad28083c9138dd3f35d9252c74fadefd184666cb8e1c7f738bd42e0

Observation 1d6f752f-6ec1-4570-a9c9-00d766f05db2 · outbound

This paper cites Visual instruction tuning, 2023 b.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Visual instruction tuning, 2023 b

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.494189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.494189Z digest=sha256:83faa99b9de2d99ea511784efe2c6788c8d41a30a5dfacbf9090288a0d8151cc

Observation c78671b4-f52e-4dc0-941e-2810b163f8bf · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.598937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.598937Z digest=sha256:d499dcd1210fd707c6d8f57d94716c29568a189fbbd15b43c908a04d49dbe61c

Observation eab265e1-d15b-4880-8499-3ffa908e199d · outbound

This paper cites Llavar: Enhanced visual instruction tuning for text-rich image understanding, 2024 b.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Llavar: Enhanced visual instruction tuning for text-rich image understanding, 2024 b

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:07.546774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:58.678265Z digest=sha256:4261d61d0eb4bcb2fce6eb5a523045c4827870f142d4cfb6c860c6d3cdf73c93

Observation a0a66828-c668-484e-9de3-7be0787a80b6 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions, 2023.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Sharegpt4v: Improving large multi-modal models with better captions, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.820250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.820250Z digest=sha256:1706494f3bf601081402a0b43fd1f60a7ed3f2fcaa8d1dcd9711ea1e7f780791

Observation 5f8b388e-d3b8-47ef-b393-3048f178eb6d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:58.895296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:58.895296Z digest=sha256:8c29457b0fec04656c87982188c4309d2043e33cb3ca76bc53fae2d94afa072c

Observation 17ff4a32-4028-4ef3-be13-82cf82910833 · outbound

This paper cites Fast r-cnn.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Fast r-cnn

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:59.109451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:59.109451Z digest=sha256:00f489371e5ca78a7379b3fbbad9076bf85de1519b2b935ab38b1457729135b9

Observation 90a65f8b-979b-4f87-a34a-38dbe01c0037 · outbound

This paper cites G-detkd: towards general distillation framework for object detectors via contrastive and semantic-guided feature imitation.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training G-detkd: towards general distillation framework for object detectors via contrastive and semantic-guided feature imitation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:07.130882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:59.257399Z digest=sha256:d8920dead22fa9db746d7f3aa22177d9f4bd3c90511a32c4e8a1532182ae11c6

Observation 659a3a35-0423-48d6-bed4-51a35d4dcdb4 · outbound

This paper cites End-to-end object detection with transformers.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training End-to-end object detection with transformers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.837984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:59.466805Z digest=sha256:8a2462ae2d428cf4aa8eb15d396a49faa98f2a8849bb0d4de81328c1324133f1

Observation 06efcf38-7a97-4ae5-bb33-b0c38f04a6d6 · outbound

This paper cites Global-local path networks for monocular depth estimation with vertical cutdepth, 2022.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Global-local path networks for monocular depth estimation with vertical cutdepth, 2022

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.645786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:59.614342Z digest=sha256:3f90dd2416cfdb4cf760aacde451e01a12a846d49137cd3fe74a444c288d9401

Observation 32f16727-5973-4eb9-8a5b-ef19c1b28cf6 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data, 2024 b.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Depth anything: Unleashing the power of large-scale unlabeled data, 2024 b

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.408486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:59.733953Z digest=sha256:3e84992643a3df00bcadf5823e740f47340139fdf4b119e4efcc71d55be664ce

Observation 703dd32b-d862-4d7e-af01-9adb8b00d82e · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:59.917204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:59.917204Z digest=sha256:73e8a6696b2ebeceee1887ad5ffa4b28672ba24617f772d02eb937d8dbc71c58

Observation d1bed38f-3a46-46f3-bd26-4df047435b01 · outbound

This paper cites Detclipv3: Towards versatile generative open-vocabulary object detection, 2024.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Detclipv3: Towards versatile generative open-vocabulary object detection, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:06.072509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:34:00.069034Z digest=sha256:d905c68cd16a3f409536d24542155a842ad06a54448a3ecea75832ca69059429

Observation 9b653fa4-3b63-43c8-aa2b-e9166cfb4795 · outbound

This paper cites Let's verify step by step.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Let's verify step by step

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.209577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.209577Z digest=sha256:aca388d9c0c6e43fa24bcdad7395c77f99925f0c20505b635411bf7099189a70

Observation 1273b247-f85f-451a-b5ce-456753c43268 · outbound

This paper cites Entropy-regularized process reward model.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Entropy-regularized process reward model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.313903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.313903Z digest=sha256:4a95a0c743d6e8612a61b9f076b75c2a850e2d49e0d7423ab9d96187f997c8fb

Observation cfe08292-f8f5-4f7a-8ffb-56766c46e2ec · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.459775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.459775Z digest=sha256:d8b9418f3bc7194aaddf287eb60b7a6f2dbebcc35a704ed43016785383e5773a

Observation 175b940e-df21-424b-bd9e-dc08f0673f5b · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.538427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.538427Z digest=sha256:c41829071f245fdcb20da043f0e7784f55e571077a9e34c95c3ba505ad119ba1

Observation e3a2a723-284e-4cfc-bbcb-7508e1d319c5 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.668832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.668832Z digest=sha256:15aa19af8ccf89b20ce08f67ed83bffede3a49916e79f4d322ab1641fbcabff0

Observation 41150ab9-a1a5-455d-a39a-e2ed661c1a3d · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:00.772326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:00.772326Z digest=sha256:16f97981e635f0c276ac368cbcafb9bd2a3cd91546db06c24e78fbcebfc50b6a

Observation 46d38e64-5e7f-4d8f-996f-544aef9cdf00 · outbound

This paper cites Reinforced self-training (rest) for language modeling.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Reinforced self-training (rest) for language modeling

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.689725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:34:00.898171Z digest=sha256:6102195ddfa8335f09d51dc7a7f2e1f4adff90e211dab38003491fd416b2891e

Observation 0f6a21d9-2b56-4249-b0d4-f50e735f049f · outbound

This paper cites Self-play fine-tuning converts weak language models to strong language models, 2024 b.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Self-play fine-tuning converts weak language models to strong language models, 2024 b

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.457636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:34:00.997086Z digest=sha256:40570ee4ef3bbf7544d10feff35b9ffc83f3a7980a12abaf1601d7e55fc62d0a

Observation 7b990001-6d13-4f8f-91b7-b7665f47f180 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Lora: Low-rank adaptation of large language models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.254752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:01.254752Z digest=sha256:dd1fae61ffdbc4246e1a72e251ba5db47afdbed7ca8374041cf5d5735dc51496

Observation 4cf8f6cb-6683-4aa0-95ff-b22f40faa089 · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.366014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:01.366014Z digest=sha256:0e5f3da6edc693d3a90473112c50dc8d1a21b6ba3cd22082c53c05a384154f3d

Observation 9f24e489-e291-46e5-b7ad-b4a4f303a15c · outbound

This paper cites RlHF-V : Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training RlHF-V : Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:05.131722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:34:01.535272Z digest=sha256:ed1f2b0abea8882ac4c9c4bd7d94691f56983a11a8657f5a1d936bda6f95a0b0

Observation 68bb5824-4ba9-4039-b49d-1727dc6a549e · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.674745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:01.674745Z digest=sha256:8f86127e2c32420fb5046e68798c835a040235ec86490c27cfadf07f20d1c48c

Observation a725c543-89c3-4ae7-8a4a-e3e1d6d57199 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:34:04.935766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:34:01.784926Z digest=sha256:db089daa442ffb04f1dbf4f6e3a393254ca8830da62f9cf9d41fc6a406854f19

Observation 497b1613-5f52-448f-88a8-1645b6bba42e · outbound

This paper cites Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.910015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:01.910015Z digest=sha256:5993796efed670dc55b5434924d699dc81e3490c75f5768132e4fc5778c2f02c

Observation 10b8cb74-dedb-4513-bbb4-17d9cd73fdbb · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.982398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:01.982398Z digest=sha256:2a3bd432c498eb3f4a4229e50f62619162c797dacccf256100d7de87d3d0da1d

Observation 964dcb89-b750-40de-87cb-9f150da3de60 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.117511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.117511Z digest=sha256:4ebc9f9aa5a566e84119f5dce2555b5af1e23e369bf0baa53683b33048371b61

Observation baff011b-f8e4-49ec-ae43-3ddf46008b7a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.276771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.276771Z digest=sha256:cfb28720caae5ebffa9070da13d6a7bb4decd8eafb4cba6ca835057a2ab630b6

Observation 54846322-0400-41c2-bddf-a584ee6cdc9c · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.392041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.392041Z digest=sha256:29079b93f08baf4d5223ce1f7d4a7bd54fe799d176e5864fcaf2d21f9c831dd1

Observation c4f294e2-d043-4c98-983d-e5fd7e3cea53 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.549287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.549287Z digest=sha256:361651ec431aaaff0b4b01f175da162eddde8e85d5b4fff4baf588cc15d0c048

Observation a18b81a1-3763-4337-ae71-4ba9b7ef5f7e · outbound

This paper cites The Llama 3 Herd of Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training The Llama 3 Herd of Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.661492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.661492Z digest=sha256:0927276bd7ca44139966e9f19bc4b3b1462f3a63a16227aa580dee0aa4b48fb5

Observation ced36032-cf5c-4953-9d69-a793bc53113a · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.816427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.816427Z digest=sha256:b3659b34957521acdc0510e42dddbfed1b595883442409812cc4f5890292ed17

Observation 210b0553-5bec-4d76-9353-5a32a79b2165 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:02.960724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:02.960724Z digest=sha256:55b40cd09d6489abcf85986f8f7c980f4ac1587d66d0eb09482be9741da51030

Observation d45607f7-0eeb-4f4a-9271-0dce9d0007ae · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:03.076564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:03.076564Z digest=sha256:9cc969a6e97d118b525da3d384568efe8b9f0e301eefc55185e3dada10332f78

Observation 456f258d-cdda-45e3-a431-c54fb8de5b8a · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:03.216795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:03.216795Z digest=sha256:97cf13889e18d0831c91696ca7402e64253561be165caac7fcb324931360754a

Observation 5876a8e5-1de7-487e-a2b7-ebbc8f195b8b · outbound

This paper cites Qwen2.5 Technical Report.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Qwen2.5 Technical Report

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:03.367306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:03.367306Z digest=sha256:f5a8ae1ac263537a0bf2e01162a9e576eadca17e245e9c1cd4b093e7792b61bd

Pith citing papers

Observation f598c56b-c547-4a9b-887c-e6ea53c849e7 · inbound

Improving Vision-language Models with Perception-centric Process Reward Models cites this paper.

Improving Vision-language Models with Perception-centric Process Reward Models VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:16.993491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T04:33:36.634359Z digest=sha256:6a558c0661f1c71142a380a5c3306fc62f0e620ef86be5ede61ef144ce726614