Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2607.09492.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 98e28a32-3613-4523-8981-671e83992518 · outbound
Multimodal Reward Hacking in Reinforcement Learning Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 989bcdf5-0de4-43b0-9795-c4cc8a4a341e · outbound
Multimodal Reward Hacking in Reinforcement Learning Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b803a01d-d71b-4ad8-9af4-bf3bdf8b73e2 · outbound
Multimodal Reward Hacking in Reinforcement Learning Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10fc225-4be8-474f-9947-e9332a7ffc41 · outbound
Multimodal Reward Hacking in Reinforcement Learning Activation Reward Models for Few-Shot Model Alignment
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81dcee44-4941-42a0-b217-466f5710df25 · outbound
Multimodal Reward Hacking in Reinforcement Learning Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a690e0b1-254d-4082-b2e0-b90008597f32 · outbound
Multimodal Reward Hacking in Reinforcement Learning Reward shaping to mitigate reward hacking in rlhf.arXiv preprint arXiv:2502.18770, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81d5aaa8-a6b5-4379-9f87-0dd63c65de62 · outbound
Multimodal Reward Hacking in Reinforcement Learning Explaining and Preventing Alignment Collapse in Iterative RLHF
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e86380e1-428a-4313-a5db-658686382711 · outbound
Multimodal Reward Hacking in Reinforcement Learning MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fab4614e-9f91-44a7-b576-41d2c481902b · outbound
Multimodal Reward Hacking in Reinforcement Learning Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c936a3e9-5eda-47b6-9c73-3cb465a1e1ab · outbound
Multimodal Reward Hacking in Reinforcement Learning Asymmetric prompt weighting for reinforcement learning with verifiable rewards.arXiv preprint arXiv:2602.11128, 2026
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13954728-6ead-44de-afac-3b708726e0e9 · outbound
Multimodal Reward Hacking in Reinforcement Learning LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70bf175c-0ac3-48b3-a38f-105623e39543 · outbound
Multimodal Reward Hacking in Reinforcement Learning Understanding reward hacking in text-to-image reinforcement learning.arXiv preprint arXiv:2601.03468, 2026
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b957c2fc-4eb6-4b63-9e60-e1234846fbef · outbound
Multimodal Reward Hacking in Reinforcement Learning VLSBench: Unveiling Visual Leakage in Multimodal Safety
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3c69c6-1b36-49f1-9d6a-e39212d688f9 · outbound
Multimodal Reward Hacking in Reinforcement Learning GPT-4o System Card
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e33ebfd3-bf91-48b1-973b-30b302a84393 · outbound
Multimodal Reward Hacking in Reinforcement Learning Do post-training algorithms actually differ? a controlled study across model scales uncovers scale-dependent ranking inversions.arXiv preprint arXiv:2603.19335, 2026
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae4585a0-8ee6-4de5-8166-64b791b8b3c7 · outbound
Multimodal Reward Hacking in Reinforcement Learning MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5325547-b6a6-435c-a3ef-e33544c56a13 · outbound
Multimodal Reward Hacking in Reinforcement Learning Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2de78b-0ff2-476e-8e7c-d91bc774fd53 · outbound
Multimodal Reward Hacking in Reinforcement Learning Towards Understanding Specification Gaming in Reasoning Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d63009a5-00f2-4d18-8539-4c9220bb4e67 · outbound
Multimodal Reward Hacking in Reinforcement Learning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 719d72d7-ac1f-485e-8b2a-94f98c9cc9d9 · outbound
Multimodal Reward Hacking in Reinforcement Learning Vlsu: Mapping the limits of joint multimodal understanding for ai safety.arXiv preprint arXiv:2510.18214, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d53086-71b7-473c-9c25-edc013723e3c · outbound
Multimodal Reward Hacking in Reinforcement Learning F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d40aca-881f-4b06-930a-cf98078d1f2c · outbound
Multimodal Reward Hacking in Reinforcement Learning Dual-bench: Measuring over-refusal and robustness in vision-language models.arXiv preprint arXiv:2510.10846, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2730f344-ed6c-46c2-9375-f2c98d0c3b49 · outbound
Multimodal Reward Hacking in Reinforcement Learning MSTS: A Multimodal Safety Test Suite for Vision-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28af7302-b8f7-4f48-885d-77815d82bb5e · outbound
Multimodal Reward Hacking in Reinforcement Learning When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 353c0b54-a1fe-4d06-adee-9740cc3fc99d · outbound
Multimodal Reward Hacking in Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17ae3cd7-bf58-4e2f-9530-8ab6beb21f43 · outbound
Multimodal Reward Hacking in Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7baf18ef-4dc7-4f3e-b099-16ab009f7235 · outbound
Multimodal Reward Hacking in Reinforcement Learning More thought, less accuracy? on the dual nature of reasoning in vision-language models.arXiv preprint arXiv:2509.25848, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb8187c-4b4f-4e6b-8b62-1bb1f0fde0a9 · outbound
Multimodal Reward Hacking in Reinforcement Learning Reward hacking as equilibrium under finite evaluation.arXiv preprint arXiv:2603.28063, 2026
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c18de61-e383-4132-824b-f5493ab66445 · outbound
Multimodal Reward Hacking in Reinforcement Learning Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33923cdc-e4b2-4088-a9a4-ca2fdce2db17 · outbound
Multimodal Reward Hacking in Reinforcement Learning Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df465da5-d28d-494a-a698-dd6bded7867e · outbound
Multimodal Reward Hacking in Reinforcement Learning Unified Reward Model for Multimodal Understanding and Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da04be5-7438-4715-a148-4cc556097ceb · outbound
Multimodal Reward Hacking in Reinforcement Learning Demystifying reinforcement learning for long-horizon tool-using agents: A comprehensive recipe.arXiv preprint arXiv:2603.21972, 2026
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4480743-638b-4363-9214-561a3a1f912c · outbound
Multimodal Reward Hacking in Reinforcement Learning Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be211758-972e-4e32-bb4f-70a2031c0cdf · outbound
Multimodal Reward Hacking in Reinforcement Learning Reward-Robust RLHF in LLMs
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419567d6-c8b2-4ee9-907b-633a9b827c4d · outbound
Multimodal Reward Hacking in Reinforcement Learning Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566278fa-5327-46a9-8c1b-24dbefcea74a · outbound
Multimodal Reward Hacking in Reinforcement Learning On the interplay of pre-training, mid-training, and rl on reasoning language models.arXiv preprint arXiv:2512.07783, 2025
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d17dbce2-9757-417a-a4be-0faecb977c04 · outbound
Multimodal Reward Hacking in Reinforcement Learning Perceptual- evidence anchored reinforced learning for multimodal reasoning.arXiv preprint arXiv:2511.18437, 2025
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b87b09-ab46-4255-ab5c-a7f29394c6ac · outbound
Multimodal Reward Hacking in Reinforcement Learning Basereward: A strong baseline for multimodal reward model.arXiv preprint arXiv:2509.16127, 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af9dee90-b69b-43f7-91f5-af013c095807 · outbound
Multimodal Reward Hacking in Reinforcement Learning MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2e385e7-cee5-4c69-be4b-fe72760b647b · outbound
Multimodal Reward Hacking in Reinforcement Learning SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b5357ef-6ce3-4e4c-ac83-faa07bdb96ad · outbound
Multimodal Reward Hacking in Reinforcement Learning Generative rlhf-v: Learning principles from multi-modal human preference.Advances in Neural Information Processing Systems, 38: 126021–126051, 2026
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c0eea38-4d5d-4bf1-9421-d4122a1bd893 · outbound
Multimodal Reward Hacking in Reinforcement Learning Omniguard: Unified safety moderation for omni-modal inputs and outputs.arXiv preprint arXiv:2512.02306, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a31e26-55e3-49f4-9243-7361b6e239b9 · outbound
Multimodal Reward Hacking in Reinforcement Learning PerPO: Perceptual Preference Optimization via Discriminative Rewarding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcb14c7c-e4f6-41f4-9f02-60eedfa91705 · outbound
Multimodal Reward Hacking in Reinforcement Learning reward hacking
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e555e2ac-a863-4f8a-9677-f1caf23279f4 · outbound
Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b05cf399-1f75-4e76-a7e4-e435aef1a43f · outbound
Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28e247b-dde4-4535-8bf2-02e73727302c · outbound
Multimodal Reward Hacking in Reinforcement Learning I cannot assist
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 064cb242-0dfd-4406-a6f8-84b808095b7a · outbound
Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1feeb64e-dc81-4782-9377-6a7085900d28 · outbound
Multimodal Reward Hacking in Reinforcement Learning label”: “Yes — No — Invalid
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba1655dd-751d-4878-b080-89785fa014d6 · outbound
Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c050b4f-3006-4fa4-8d8e-5e6cfef03e44 · outbound
Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eaeca36-6819-4e41-a019-bbb477199133 · outbound
Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f5d96d-5fd2-4c19-9a49-a66983734f14 · outbound
Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be453a2d-5fed-4e16-9134-5838b9456212 · outbound
Multimodal Reward Hacking in Reinforcement Learning sft score
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeff1747-e8a9-4de2-b986-59d9dc8b7d03 · outbound
Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc2ca65-285d-4adf-adf6-73c70c993a70 · outbound
Multimodal Reward Hacking in Reinforcement Learning I cannot assist. The image contains knives
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb47c47a-30e2-4c5d-b2e1-8a47747ea30c · outbound
Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a10d9a-3a22-4452-8e9f-82be29fc6eff · outbound
Multimodal Reward Hacking in Reinforcement Learning I cannot help with this request because the image shows instructions for making a weapon, which I cannot assist with
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.