Pith. sign in

Paper Citation Record · LEDGER

Generative RLHF-V: Learning Principles from Multi-modal Human Preference

As of 8 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 6 inbound Pith citation observations for arXiv:2505.18531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18531 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:56.530687Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:46:36.081851Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.415184Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved44
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52a99462-90e1-4ab7-bd0c-14e5fe19f776 · outbound

This paper cites Machine behaviour.Nature, 568(7753):477–486, 2019.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Machine behaviour.Nature, 568(7753):477–486, 2019

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:59.349875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:34:51.781741Z digest=sha256:a857241eb11e3aafe1feeaeb12e832f99b23f435a87c1bbbb7e144ad7b38b97e

Observation 1359e8a1-2d26-4219-a312-5d91c55f3c93 · outbound

This paper cites Position: The platonic representation hypothesis.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Position: The platonic representation hypothesis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:59.063038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:34:51.856669Z digest=sha256:0bcfb5a51182671d59a99f28b1ad3eacefb628f1e319780580d77dc5f0c4cd9f

Observation e3134af0-548e-4fd7-baa0-593b445b2844 · outbound

This paper cites PaLM 2 Technical Report.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference PaLM 2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:51.982149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:51.982149Z digest=sha256:f001493e1b508a8834887b6f950a4597187db6573b6c77ea890560aa7fc57337

Observation cd735611-0f77-43ff-9c69-1e1ee27c4c29 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.093446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.093446Z digest=sha256:225029b4937fffd6efc95705d317a4f4146c13b1c58a55d7843df27709029733

Observation 623e31c2-459d-446f-ada9-acb8138d1d65 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.213552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.213552Z digest=sha256:af059416d9d91b4d3605244b5d474938b955526846e672989d2c6e7b970da114

Observation f8f34cbe-3bb1-4204-aef1-f334d06d5fc7 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.342099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.342099Z digest=sha256:30e462ed0ad7634321a2b25eedc07665c501958ba3b52dec0a08004623bb123a

Observation e6d81959-21c7-44ba-b9e0-e5d89af0b7ee · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.484882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.484882Z digest=sha256:dea6d4db15841700bdbaec06cd34f168eb1050728638cd31b79176fd2099d50f

Observation d10f924b-abf7-43d2-bd37-e92e23c18266 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.613268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.613268Z digest=sha256:bf56af5cf8029fc4785a99e400c2c00cfa496064573d61b977d49d5ce8e5fc66

Observation d727a3bb-bf2b-46f0-a4e8-897b4c468adc · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.685004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.685004Z digest=sha256:58852cc4fdfc86688ca5b4b29ed458a399e6d48d0cbd74bc5beb80eb8cc8a2d4

Observation 8ee10568-f094-4b98-9881-81503bf2fc48 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.745966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.745966Z digest=sha256:3608cdc99d139cbe6e9f2baabb4f1f6cbd0b0899a541ed6b9af9ca8497d2cb27

Observation b6922c7f-d46c-4afa-8924-5804c5de8353 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.837437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.837437Z digest=sha256:b5143837bd062f154d4850c7f22e541394efe1aaffd31cff846080c771f72f38

Observation 29e48177-1199-416e-96af-85929b9fff2f · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference AI Alignment: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.908750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.908750Z digest=sha256:8e3ab6f67a9bbfd8bdfcbd12e3c5cbc9c9b2a93c0d84cb0e8f95529516082fb2

Observation 6d919e31-cbdf-4830-bb6e-b5f1ee4e564a · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference A General Language Assistant as a Laboratory for Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.995877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.995877Z digest=sha256:938d07afcca447ac5b65f787c5b77fe14b2889122969636cb5c0637f95402848

Observation dfc96d4b-a586-44b6-89c8-d0f0201f66e8 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.070027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.070027Z digest=sha256:ac54a753df541bec3c5501c41bb028c84cdd54214a1fbb4446e8842cd5067524

Observation ce052be3-c027-4d9c-8e25-20d10f7eb5df · outbound

This paper cites MM-IFEngine: Towards Multimodal Instruction Following.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference MM-IFEngine: Towards Multimodal Instruction Following

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.141416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.141416Z digest=sha256:f1bd85952d2e2f5d06fb8d4ba527a46fe0a01b13d8b041ffe0f82c9c209379dd

Observation b88ba47c-6566-419d-8439-1d04142bb021 · outbound

This paper cites GPT-4 Technical Report.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.195079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.195079Z digest=sha256:2b03fadf504c1afdd931855c4d472f1ea784d9ac75db0fe962c997d13a8724ab

Observation 935d3d34-b17a-4890-b3d3-d729b735bc58 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Safe rlhf: Safe reinforcement learning from human feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:58.796943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:34:53.248426Z digest=sha256:777136bfa4ce029c4e795b40cf6dd2e888534b0f5771e14717f30c57d55b14a8

Observation 9ab55e76-ce5c-49ec-99b3-912d3ec48425 · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Aligning large multimodal models with factually augmented rlhf

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:58.467639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:34:53.310333Z digest=sha256:53756403dc2f44ffea0646d74e0bfdb934de15d183f9dd474c3c1cacf55926a7

Observation d1dc4680-32bd-420c-887c-cec55f5194ae · outbound

This paper cites Scaling laws for reward model overoptimization.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Scaling laws for reward model overoptimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.381691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.381691Z digest=sha256:d14b0fe81a310a2c94d4fb2e8e7782c0fdab732d5085a15b61ed5ba9b6ca71c7

Observation a78ac211-9a6f-4e48-8125-5974f0f11af6 · outbound

This paper cites Sequence to sequence reward modeling: Improving rlhf by language feedback.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Sequence to sequence reward modeling: Improving rlhf by language feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.463689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.463689Z digest=sha256:ca475d68b2e635a0524796f24b8277eb319daf60f14a9564ac7aee883aaf0231

Observation 053d29f9-a696-4f8b-bf2c-3e99b8a48106 · outbound

This paper cites Critique-out-Loud Reward Models.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Critique-out-Loud Reward Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.519748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.519748Z digest=sha256:71579743a4cacebf0f2ee8da21c29c292f1ea2da2dfa374e785ba52b0c996772

Observation 3e36d3c2-5495-4f18-8ed3-08938b55deb7 · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Generative verifiers: Reward modeling as next-token prediction

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:58.190130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:34:53.612089Z digest=sha256:6aa402258093fa9658248b4264155e5fe0d02d4cfe99e7d480f43e9dc0f773ea

Observation 1763faa3-2891-4513-aeef-99914871d801 · outbound

This paper cites Generative Reward Models.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Generative Reward Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.738399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.738399Z digest=sha256:39052ac549f5240144f83687c346677907d1ba3e2542db61551bd762025ca3e6

Observation 4fa3500f-5b2d-4877-8ebe-ebbdf249ef2b · outbound

This paper cites Unified Reward Model for Multimodal Understanding and Generation.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Unified Reward Model for Multimodal Understanding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.833434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.833434Z digest=sha256:734b52b6a2a82a2383437c5babf3b551a08e8c6eb294383c7d145f4f700de4f8

Observation a30d78d2-7615-43c0-8935-7779e8671eeb · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.906318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.906318Z digest=sha256:ce7a4862d06e0c25b0b8190f3c5d612cbd7d2433afd6e78f148ff3bb9c597a33

Observation 80c05f73-3ac8-4538-ba17-7cd5668d1964 · outbound

This paper cites Inference-time scaling for generalist reward modeling, 2025.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Inference-time scaling for generalist reward modeling, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.000417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.000417Z digest=sha256:8664787a5b60f4c790491cc1c614cd089f8cfc5e19fa421d0561b71ee2c94c41

Observation cef6e9f4-d236-4ce6-8e38-760a869cdd3b · outbound

This paper cites A survey of multimodel large language models.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference A survey of multimodel large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.081292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.081292Z digest=sha256:619eb4c5d304015a4f75a6875b5e502ab48440ee3e9437d3291ca2699658424a

Observation ee03b3fa-d2de-4e87-a295-f23b14920c33 · outbound

This paper cites A Survey on Progress in LLM Alignment from the Perspective of Reward Design.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference A Survey on Progress in LLM Alignment from the Perspective of Reward Design

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.149826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.149826Z digest=sha256:0e5d93a4adfe4eb3fb8ec4639bc8ab841a08c06dee208122e78828acb9654777

Observation fe590e6b-b6e5-4f90-a71c-5e6eb9fd070e · outbound

This paper cites Beyond Scalar Reward Model: Learning Generative Judge from Preference Data.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.230084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.230084Z digest=sha256:f87074948faeaa8547a1394ef6b2820633b468b8c8a81c5d359189de53248cbb

Observation a4e55663-8b76-49a3-ad9d-a4f21805e668 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:57.853980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:34:54.304404Z digest=sha256:34c2657ae06db80f3f3b3b8800dd1fe898f8e4c6d067b6482a5cb0ab489d77b1

Observation 9d24378c-c3ff-4b89-ac05-ab2d8925661b · outbound

This paper cites Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.380007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.380007Z digest=sha256:8a0a980dc99f97bc54374ed7f5f17c2ef4fb6339b6a973c7297c148ce01041d1

Observation 79be537c-ece5-4495-b294-07b595d9fd67 · outbound

This paper cites A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.497908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.497908Z digest=sha256:d19efd9200d721aa178bb6a6a6d99f5ff2454c46d6bb8bb06d909d82c2a00f40

Observation 5d07aa45-5ab5-41ed-b037-578289eacc1e · outbound

This paper cites Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:57.634491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:34:54.578435Z digest=sha256:05a2b4155681a6606df54d9b7cab7ef0bd2fe512daad53572f30ac066e03775f

Observation d1919462-3483-4a39-b78c-d3c2f717c9c6 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.662351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.662351Z digest=sha256:294dc1f382c7a772465d3a0817ca66501a5526bc75dfa51d26d842c22e135ce7

Observation 6cdd35f8-35e6-4c18-9d7c-3b05baadb601 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Constitutional AI: Harmlessness from AI Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.754538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.754538Z digest=sha256:bb791a3fe8b3b4a9f5e3dc89e1ba2a5dc3d7c4169d8830c4f10b749a89033d0f

Observation 7395b120-d87a-408a-94b0-37364b8e45e3 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Direct preference optimization: Your language model is secretly a reward model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.885470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.885470Z digest=sha256:f87348b5e40134d07f03241c97a45f93818664eb4334243d55ce0ca25caf0997

Observation cc82f5a6-d743-41f8-be7d-16ad5d48d062 · outbound

This paper cites CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.013457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.013457Z digest=sha256:f0453c7b8117ff98f8b1ce732906c5113d6bd486f992c0cd30625f1ce329fb1b

Observation 62de62e2-b0ec-4e6f-8e41-c84fe1f9368e · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.112648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.112648Z digest=sha256:2a413b9b937b526b1a02aac161c2fe9315633a9236bc92e930f5e3aa60bd1b51

Observation 2e2966af-922c-4976-8a13-ebd3c350c01d · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.184677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.184677Z digest=sha256:eba095bae6c8d6b65cda045cccec7d8b23f634825326cce90737f7d8b64b0084

Observation ae56dbab-7a2e-43bd-8dcc-caf3f8bb724f · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.253963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.253963Z digest=sha256:d84640b8d7d23f627805b04470961edbffefd042503d29241a31cdab863515e0

Observation 1b496ee1-faab-472b-86ff-85473e424dda · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.364341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.364341Z digest=sha256:bf7a28889d0d989eebf0c413b8c92ec8e7c4b3bd1987b060cc199c735b589cbe

Observation 3db04bc1-b713-4df1-a136-bbdcad8cefd3 · outbound

This paper cites Qwen2.5-VL Technical Report.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Qwen2.5-VL Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.460092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.460092Z digest=sha256:7321ebf666a2e4cf4dad89be26f5d6f7fe2cdf96a96be5e52965de6139cfb817

Observation e4fbbf22-fdf0-409d-b2e1-5dc8c0824026 · outbound

This paper cites Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.554751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.554751Z digest=sha256:ced91eef61b7bea36eb162c6e301113e2156d1ca8e1e15b8642b524291a5e328

Observation 418727c9-566c-4acd-87f1-fd81820d4ac0 · outbound

This paper cites Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.643156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.643156Z digest=sha256:95ed8e1108fcbd51d0bcdfef9aa460713c218332c97b8cfc6846285cca7a635d

Observation b6ad1ebc-e9b2-4191-9db1-a3959039a924 · outbound

This paper cites MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.736257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.736257Z digest=sha256:dd9b4b95050e256f1104e3ee9298bd51414a32f6f57495d4b1a21abb463108bd

Observation 86cffb9d-9813-42bf-9d7e-1f387ca7258c · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.819660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.819660Z digest=sha256:7cee051ac01c14eb9258074ec8da39d18fc96b4d6b8445d66f74019a0cafbed3

Observation 9941a29e-9d71-4f62-9880-7454f21155b3 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.904276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.904276Z digest=sha256:3191ed20a2a9f64c9b13cb4a41ac8ea03cb97d450d6ba07cb4384d3f4b284fe0

Observation b98aa5a5-b8eb-4955-b336-a41c1b9fd295 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:57.467100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:34:56.006470Z digest=sha256:54aaa3eb2397424f5b29c9d4d77228c81ee7f9f14a318dd86223fc672400ff7c

Observation e2d2eab0-142b-433a-bce1-7e25feec12ab · outbound

This paper cites MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:56.100138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:56.100138Z digest=sha256:85ab291cc8400fd06adb6c812e0eab11e4a9bee3a618f5d3da8f2c867ffcfe4f

Observation ed1870d8-1ef6-4e2d-9138-067accb5dffb · outbound

This paper cites Mm-safetybench: A benchmark for safety evaluation of multimodal large language models.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Mm-safetybench: A benchmark for safety evaluation of multimodal large language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:56.195672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:56.195672Z digest=sha256:665f5f9fc7fc4f2e171ef8e734e1a103b967f52880cb0005f70f9a5f22142049

Observation 136e2999-40e0-4b35-bac4-99a0026ac9fe · outbound

This paper cites Multimodal Situational Safety.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Multimodal Situational Safety

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:56.289308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:56.289308Z digest=sha256:72da52bfc87542079e1cfb67af7abb9fea109ae087dc79e8f459d4c2ed5119a9

Observation e700ad34-5f58-49d1-ae35-a126ddc9674c · outbound

This paper cites Concrete Problems in AI Safety.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Concrete Problems in AI Safety

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:56.359753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:56.359753Z digest=sha256:bbaf563e84a53e224e3f3734f92102c5446a7d8911d8c9acfc8ad3a58ee4fc6e

Observation 5e4a659f-7d3a-43f7-a589-0ffff9a3f635 · outbound

This paper cites The effects of reward misspecification: Mapping and mitigating misaligned models.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference The effects of reward misspecification: Mapping and mitigating misaligned models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:57.320747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:34:56.459348Z digest=sha256:1359bfaf3543f837fa310aad76f45ec6ff59e5b1dd38b946d016d2152dba756b

Observation ec594816-cff5-4048-a875-53614148ed77 · outbound

This paper cites image-text sequence understanding.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference image-text sequence understanding

Reference 54

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:34:57.123628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:34:56.530687Z digest=sha256:a940da590a628717fc34b8b7d16b8e41bc2e53b9ce7399d7f85ca444cf09e8fb

Pith citing papers

Observation 728a9f85-9465-4cee-b8ae-81837063aed8 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning Generative RLHF-V: Learning Principles from Multi-modal Human Preference

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:32:22.432616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:aee2f4a7b858d7ff3bceb7b3aad88b8d27107240df0da00d7202621f7fc06222

Observation 82156c10-6644-4c2e-9a4a-327d31728a7b · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Generative RLHF-V: Learning Principles from Multi-modal Human Preference

Reference 184

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.711508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:1b1d1dd0b9e72d53c1e5228ae66b2ec2b2e59a8de5df27ceaf8dea96450d30f0

Observation 5da103f6-4edf-46d2-9d1c-883a3a2601ec · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Generative RLHF-V: Learning Principles from Multi-modal Human Preference

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:04.075359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:dec1b8eb10a25bd29a265ae6481879c89b7369945a59f4718dc91b7136f932f3

Observation 984b7b75-dd33-46db-9b93-a28a21cc54ed · inbound

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation cites this paper.

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation Generative RLHF-V: Learning Principles from Multi-modal Human Preference

Reference 290

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:56.761079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T01:46:36.081851Z digest=sha256:06e7c3fc9c5745479a25177ef494ed0d6c25b391ed389f457f4a75aaaa15a8a9

Observation 674ecc18-d9f2-4006-87c0-b9d1b71c5b2f · inbound

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning? cites this paper.

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning? Generative RLHF-V: Learning Principles from Multi-modal Human Preference

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:37:14.732384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T21:56:25.350827Z digest=sha256:199637ddd312710f39423c6014eafbee92fda68eda86064557035c60353a4c21

Observation d3cc7c17-880b-4070-bf1b-06e0208dc4e8 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Generative RLHF-V: Learning Principles from Multi-modal Human Preference

Reference 277

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.416476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:aa69ba4571e58a624e36cf780f888b728356d649bf20c487148d4c74cb0050f3