Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:56.530687Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 6 inbound Pith citation observations for arXiv:2505.18531.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:56.530687Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T01:46:36.081851Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T01:37:30.415184Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 52a99462-90e1-4ab7-bd0c-14e5fe19f776 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Machine behaviour.Nature, 568(7753):477–486, 2019
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1359e8a1-2d26-4219-a312-5d91c55f3c93 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Position: The platonic representation hypothesis
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e3134af0-548e-4fd7-baa0-593b445b2844 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference PaLM 2 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd735611-0f77-43ff-9c69-1e1ee27c4c29 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623e31c2-459d-446f-ada9-acb8138d1d65 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8f34cbe-3bb1-4204-aef1-f334d06d5fc7 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6d81959-21c7-44ba-b9e0-e5d89af0b7ee · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d10f924b-abf7-43d2-bd37-e92e23c18266 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d727a3bb-bf2b-46f0-a4e8-897b4c468adc · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee10568-f094-4b98-9881-81503bf2fc48 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6922c7f-d46c-4afa-8924-5804c5de8353 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29e48177-1199-416e-96af-85929b9fff2f · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference AI Alignment: A Comprehensive Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d919e31-cbdf-4830-bb6e-b5f1ee4e564a · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference A General Language Assistant as a Laboratory for Alignment
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfc96d4b-a586-44b6-89c8-d0f0201f66e8 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce052be3-c027-4d9c-8e25-20d10f7eb5df · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference MM-IFEngine: Towards Multimodal Instruction Following
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88ba47c-6566-419d-8439-1d04142bb021 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference GPT-4 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 935d3d34-b17a-4890-b3d3-d729b735bc58 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Safe rlhf: Safe reinforcement learning from human feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ab55e76-ce5c-49ec-99b3-912d3ec48425 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Aligning large multimodal models with factually augmented rlhf
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1dc4680-32bd-420c-887c-cec55f5194ae · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Scaling laws for reward model overoptimization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a78ac211-9a6f-4e48-8125-5974f0f11af6 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Sequence to sequence reward modeling: Improving rlhf by language feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 053d29f9-a696-4f8b-bf2c-3e99b8a48106 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Critique-out-Loud Reward Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e36d3c2-5495-4f18-8ed3-08938b55deb7 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Generative verifiers: Reward modeling as next-token prediction
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1763faa3-2891-4513-aeef-99914871d801 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Generative Reward Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fa3500f-5b2d-4877-8ebe-ebbdf249ef2b · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Unified Reward Model for Multimodal Understanding and Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a30d78d2-7615-43c0-8935-7779e8671eeb · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference LLaVA-Critic: Learning to Evaluate Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c05f73-3ac8-4538-ba17-7cd5668d1964 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Inference-time scaling for generalist reward modeling, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cef6e9f4-d236-4ce6-8e38-760a869cdd3b · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference A survey of multimodel large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee03b3fa-d2de-4e87-a295-f23b14920c33 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference A Survey on Progress in LLM Alignment from the Perspective of Reward Design
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe590e6b-b6e5-4f90-a71c-5e6eb9fd070e · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4e55663-8b76-49a3-ad9d-a4f21805e668 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d24378c-c3ff-4b89-ac05-ab2d8925661b · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79be537c-ece5-4495-b294-07b595d9fd67 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d07aa45-5ab5-41ed-b037-578289eacc1e · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1919462-3483-4a39-b78c-d3c2f717c9c6 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cdd35f8-35e6-4c18-9d7c-3b05baadb601 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Constitutional AI: Harmlessness from AI Feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7395b120-d87a-408a-94b0-37364b8e45e3 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Direct preference optimization: Your language model is secretly a reward model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc82f5a6-d743-41f8-be7d-16ad5d48d062 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62de62e2-b0ec-4e6f-8e41-c84fe1f9368e · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e2966af-922c-4976-8a13-ebd3c350c01d · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae56dbab-7a2e-43bd-8dcc-caf3f8bb724f · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b496ee1-faab-472b-86ff-85473e424dda · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db04bc1-b713-4df1-a136-bbdcad8cefd3 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Qwen2.5-VL Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4fbbf22-fdf0-409d-b2e1-5dc8c0824026 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 418727c9-566c-4acd-87f1-fd81820d4ac0 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6ad1ebc-e9b2-4191-9db1-a3959039a924 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86cffb9d-9813-42bf-9d7e-1f387ca7258c · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9941a29e-9d71-4f62-9880-7454f21155b3 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98aa5a5-b8eb-4955-b336-a41c1b9fd295 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Mm-vet: Evaluating large multimodal models for integrated capabilities
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2d2eab0-142b-433a-bce1-7e25feec12ab · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1870d8-1ef6-4e2d-9138-067accb5dffb · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Mm-safetybench: A benchmark for safety evaluation of multimodal large language models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 136e2999-40e0-4b35-bac4-99a0026ac9fe · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Multimodal Situational Safety
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e700ad34-5f58-49d1-ae35-a126ddc9674c · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Concrete Problems in AI Safety
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e4a659f-7d3a-43f7-a589-0ffff9a3f635 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference The effects of reward misspecification: Mapping and mitigating misaligned models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec594816-cff5-4048-a875-53614148ed77 · outbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference image-text sequence understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 728a9f85-9465-4cee-b8ae-81837063aed8 · inbound
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning Generative RLHF-V: Learning Principles from Multi-modal Human Preference
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 82156c10-6644-4c2e-9a4a-327d31728a7b · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Generative RLHF-V: Learning Principles from Multi-modal Human Preference
Reference 184
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5da103f6-4edf-46d2-9d1c-883a3a2601ec · inbound
DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Generative RLHF-V: Learning Principles from Multi-modal Human Preference
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 984b7b75-dd33-46db-9b93-a28a21cc54ed · inbound
Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation Generative RLHF-V: Learning Principles from Multi-modal Human Preference
Reference 290
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 674ecc18-d9f2-4006-87c0-b9d1b71c5b2f · inbound
VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning? Generative RLHF-V: Learning Principles from Multi-modal Human Preference
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3cc7c17-880b-4070-bf1b-06e0208dc4e8 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Generative RLHF-V: Learning Principles from Multi-modal Human Preference
Reference 277
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.