Pith. sign in

Paper Citation Record · LEDGER

Multimodal Reward Hacking in Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2607.09492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09492 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98e28a32-3613-4523-8981-671e83992518 · outbound

This paper cites Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards.

Multimodal Reward Hacking in Reinforcement Learning Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:bb22053753a3999a7de5546c2564d6b0ccfdb87f8ed1ec87beaf543bcd3e9317

Observation 989bcdf5-0de4-43b0-9795-c4cc8a4a341e · outbound

This paper cites Qwen3-VL Technical Report.

Multimodal Reward Hacking in Reinforcement Learning Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:dbaa4d86ccd8b964bcbe49414d647ff9894763b89bfd5aec227648cc16dea77d

Observation b803a01d-d71b-4ad8-9af4-bf3bdf8b73e2 · outbound

This paper cites Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes.

Multimodal Reward Hacking in Reinforcement Learning Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:d496bcaaa32aaf213f36998aef201f5ed18076f563d62a7e06d1008a97a01a4d

Observation e10fc225-4be8-474f-9947-e9332a7ffc41 · outbound

This paper cites Activation Reward Models for Few-Shot Model Alignment.

Multimodal Reward Hacking in Reinforcement Learning Activation Reward Models for Few-Shot Model Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:ea267e612d41090b277ddb3feba07b1178570833ab81ab71e0509de5ea9ed5f1

Observation 81dcee44-4941-42a0-b217-466f5710df25 · outbound

This paper cites Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling.

Multimodal Reward Hacking in Reinforcement Learning Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:112565f1488cf1db89aee8c000bbabd0fd8463fd3be1a5ccf3a42c892fa01405

Observation a690e0b1-254d-4082-b2e0-b90008597f32 · outbound

This paper cites Reward shaping to mitigate reward hacking in rlhf.arXiv preprint arXiv:2502.18770, 2025.

Multimodal Reward Hacking in Reinforcement Learning Reward shaping to mitigate reward hacking in rlhf.arXiv preprint arXiv:2502.18770, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:75c3ab1ed1ec2589568a044f9f60d485249ecf5d723a177df8dd614f49fa908e

Observation 81d5aaa8-a6b5-4379-9f87-0dd63c65de62 · outbound

This paper cites Explaining and Preventing Alignment Collapse in Iterative RLHF.

Multimodal Reward Hacking in Reinforcement Learning Explaining and Preventing Alignment Collapse in Iterative RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:50134abe4a6c6e29fd2f9ce56e51ff67aa17789caafefbbf41957f7e7f1d1692

Observation e86380e1-428a-4313-a5db-658686382711 · outbound

This paper cites MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models.

Multimodal Reward Hacking in Reinforcement Learning MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:f181230a6011c6cdd1f28337311e903e4543a0b6734289b158226b8ee3110feb

Observation fab4614e-9f91-44a7-b576-41d2c481902b · outbound

This paper cites Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.

Multimodal Reward Hacking in Reinforcement Learning Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:858b0a1a600ae4d477d4fe495796fa83d8639064354772ed97ec760412a2e729

Observation c936a3e9-5eda-47b6-9c73-3cb465a1e1ab · outbound

This paper cites Asymmetric prompt weighting for reinforcement learning with verifiable rewards.arXiv preprint arXiv:2602.11128, 2026.

Multimodal Reward Hacking in Reinforcement Learning Asymmetric prompt weighting for reinforcement learning with verifiable rewards.arXiv preprint arXiv:2602.11128, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:9ef01c164b282a7950e6ee1e8636fd639c79c57a957b6f8601e0a4c1178f56eb

Observation 13954728-6ead-44de-afac-3b708726e0e9 · outbound

This paper cites LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models.

Multimodal Reward Hacking in Reinforcement Learning LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:f18e39061beaae72e2844e6ca28373cbecff79c5ee0211509a1bfeae2b87ad86

Observation 70bf175c-0ac3-48b3-a38f-105623e39543 · outbound

This paper cites Understanding reward hacking in text-to-image reinforcement learning.arXiv preprint arXiv:2601.03468, 2026.

Multimodal Reward Hacking in Reinforcement Learning Understanding reward hacking in text-to-image reinforcement learning.arXiv preprint arXiv:2601.03468, 2026

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:c114d49902fa6c338d6153ce600cb0dee19f65f8e13fa8ba6c31e3b9d3db4090

Observation b957c2fc-4eb6-4b63-9e60-e1234846fbef · outbound

This paper cites VLSBench: Unveiling Visual Leakage in Multimodal Safety.

Multimodal Reward Hacking in Reinforcement Learning VLSBench: Unveiling Visual Leakage in Multimodal Safety

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:1994f4906b766d4231a438928c59e334725c60df433a7be9b45eb1f8a46a6942

Observation af3c69c6-1b36-49f1-9d6a-e39212d688f9 · outbound

This paper cites GPT-4o System Card.

Multimodal Reward Hacking in Reinforcement Learning GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:ddcc30ce9778d32ff562710981e225d397c6d0cbec61478d2493ae861f2a045f

Observation e33ebfd3-bf91-48b1-973b-30b302a84393 · outbound

This paper cites Do post-training algorithms actually differ? a controlled study across model scales uncovers scale-dependent ranking inversions.arXiv preprint arXiv:2603.19335, 2026.

Multimodal Reward Hacking in Reinforcement Learning Do post-training algorithms actually differ? a controlled study across model scales uncovers scale-dependent ranking inversions.arXiv preprint arXiv:2603.19335, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:c6dd858b0339d0968ede67d87c35c86cb818c08c95c151edb26c84e94973faa1

Observation ae4585a0-8ee6-4de5-8166-64b791b8b3c7 · outbound

This paper cites MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models.

Multimodal Reward Hacking in Reinforcement Learning MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:8da34380f6b3c6bc317f428b9cfd50fa2f6070352449344f24a1ab45b55fc8d0

Observation e5325547-b6a6-435c-a3ef-e33544c56a13 · outbound

This paper cites Robust Optimization for Mitigating Reward Hacking with Correlated Proxies.

Multimodal Reward Hacking in Reinforcement Learning Robust Optimization for Mitigating Reward Hacking with Correlated Proxies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:201afc5aab728f7eb36e85ee9c77c8904c872e99007f802124035e088cf6630b

Observation 1c2de78b-0ff2-476e-8e7c-d91bc774fd53 · outbound

This paper cites Towards Understanding Specification Gaming in Reasoning Models.

Multimodal Reward Hacking in Reinforcement Learning Towards Understanding Specification Gaming in Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:bb1285193ecb52c407ee4f8b7ca47d0e8b55d67b84ecdee376f846c8342cd511

Observation d63009a5-00f2-4d18-8539-4c9220bb4e67 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Multimodal Reward Hacking in Reinforcement Learning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:a90544ecc08f5bdda70ecfc33f6c8950e552c15c6713b27f39e7dfef80f96e37

Observation 719d72d7-ac1f-485e-8b2a-94f98c9cc9d9 · outbound

This paper cites Vlsu: Mapping the limits of joint multimodal understanding for ai safety.arXiv preprint arXiv:2510.18214, 2025.

Multimodal Reward Hacking in Reinforcement Learning Vlsu: Mapping the limits of joint multimodal understanding for ai safety.arXiv preprint arXiv:2510.18214, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:6cc44bd7170929cb6f6fc90c5628490e6ccf9705d1b85eb778403d7113280811

Observation 54d53086-71b7-473c-9c25-edc013723e3c · outbound

This paper cites F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare.

Multimodal Reward Hacking in Reinforcement Learning F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:189d98e65fc37d5099a01a13254a446498c2918aeb860f8b3a1d2e41c18d0743

Observation 85d40aca-881f-4b06-930a-cf98078d1f2c · outbound

This paper cites Dual-bench: Measuring over-refusal and robustness in vision-language models.arXiv preprint arXiv:2510.10846, 2025.

Multimodal Reward Hacking in Reinforcement Learning Dual-bench: Measuring over-refusal and robustness in vision-language models.arXiv preprint arXiv:2510.10846, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:6da8078deb59df1638c65875fb8f9ac4f51554c1fead563b0440e23dcfca46d0

Observation 2730f344-ed6c-46c2-9375-f2c98d0c3b49 · outbound

This paper cites MSTS: A Multimodal Safety Test Suite for Vision-Language Models.

Multimodal Reward Hacking in Reinforcement Learning MSTS: A Multimodal Safety Test Suite for Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:671cf77c3210b23659dd12858089929caa8ec7de6a68dfbb677537aa5ff4f637

Observation 28af7302-b8f7-4f48-885d-77815d82bb5e · outbound

This paper cites When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient.

Multimodal Reward Hacking in Reinforcement Learning When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:4a6708f067f28149969f111c3919a00a662be226a15590e18dd12d7840c71d76

Observation 353c0b54-a1fe-4d06-adee-9740cc3fc99d · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Multimodal Reward Hacking in Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:bd8282f34395ebb554d89b7d11bfc4d09f6489a0adf3d05b1cfb311279a68f74

Observation 17ae3cd7-bf58-4e2f-9530-8ab6beb21f43 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Multimodal Reward Hacking in Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:92ccfc8ed3343aa6709f5bb7394cfa3279affe99b7bb6706f9732cdc57db8140

Observation 7baf18ef-4dc7-4f3e-b099-16ab009f7235 · outbound

This paper cites More thought, less accuracy? on the dual nature of reasoning in vision-language models.arXiv preprint arXiv:2509.25848, 2025.

Multimodal Reward Hacking in Reinforcement Learning More thought, less accuracy? on the dual nature of reasoning in vision-language models.arXiv preprint arXiv:2509.25848, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:d7d62e2389937528a7393ee9e6986bac927ef9fbb00e65ecd060ac461238e03b

Observation feb8187c-4b4f-4e6b-8b62-1bb1f0fde0a9 · outbound

This paper cites Reward hacking as equilibrium under finite evaluation.arXiv preprint arXiv:2603.28063, 2026.

Multimodal Reward Hacking in Reinforcement Learning Reward hacking as equilibrium under finite evaluation.arXiv preprint arXiv:2603.28063, 2026

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:262095bf9a8a4bfc38b7fd9c1c2a35447f38de7dd58885b82cbfe0388a08fe0c

Observation 2c18de61-e383-4132-824b-f5493ab66445 · outbound

This paper cites Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges.

Multimodal Reward Hacking in Reinforcement Learning Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:f946c9a0af868d43c78638ea3b735a90165b6d3abc05061e191d25dedccc8dbc

Observation 33923cdc-e4b2-4088-a9a4-ca2fdce2db17 · outbound

This paper cites Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning.

Multimodal Reward Hacking in Reinforcement Learning Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:2aa839987b6562c9b491ebbd2d40530088a31de40e17f4c1fd0545c45a3bae7d

Observation df465da5-d28d-494a-a698-dd6bded7867e · outbound

This paper cites Unified Reward Model for Multimodal Understanding and Generation.

Multimodal Reward Hacking in Reinforcement Learning Unified Reward Model for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:32f87711f152045fc148d7f64809b8b94a007dc454de1b6f8d1e30c414478355

Observation 4da04be5-7438-4715-a148-4cc556097ceb · outbound

This paper cites Demystifying reinforcement learning for long-horizon tool-using agents: A comprehensive recipe.arXiv preprint arXiv:2603.21972, 2026.

Multimodal Reward Hacking in Reinforcement Learning Demystifying reinforcement learning for long-horizon tool-using agents: A comprehensive recipe.arXiv preprint arXiv:2603.21972, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:a6cd27e58f05417e48c267f5a3a7e96e3bd6f56092a6d6473797f924b1715e33

Observation a4480743-638b-4363-9214-561a3a1f912c · outbound

This paper cites Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients.

Multimodal Reward Hacking in Reinforcement Learning Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:4140cf7b46356d67d83bb2fbb8c3a543410c0d2039eddea709cec3b507ff169d

Observation be211758-972e-4e32-bb4f-70a2031c0cdf · outbound

This paper cites Reward-Robust RLHF in LLMs.

Multimodal Reward Hacking in Reinforcement Learning Reward-Robust RLHF in LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:1a86db488f1d62cfcbe6628ea72e477450751a15c723abc4bc43d70c5a6c9d76

Observation 419567d6-c8b2-4ee9-907b-633a9b827c4d · outbound

This paper cites Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO.

Multimodal Reward Hacking in Reinforcement Learning Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:547233c17d6aa35b06dcb5726d0e1ded420b273d19c2550f9401f5e27bab9b19

Observation 566278fa-5327-46a9-8c1b-24dbefcea74a · outbound

This paper cites On the interplay of pre-training, mid-training, and rl on reasoning language models.arXiv preprint arXiv:2512.07783, 2025.

Multimodal Reward Hacking in Reinforcement Learning On the interplay of pre-training, mid-training, and rl on reasoning language models.arXiv preprint arXiv:2512.07783, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:4b490c202c53fe877485bceb59e1448c02a308393634a2b3ebeec1bfc427e809

Observation d17dbce2-9757-417a-a4be-0faecb977c04 · outbound

This paper cites Perceptual- evidence anchored reinforced learning for multimodal reasoning.arXiv preprint arXiv:2511.18437, 2025.

Multimodal Reward Hacking in Reinforcement Learning Perceptual- evidence anchored reinforced learning for multimodal reasoning.arXiv preprint arXiv:2511.18437, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:53a0a8ac4ed1eb1111ecd8e24b4b0cf1df5d8896be7a78154832f655406181e6

Observation c8b87b09-ab46-4255-ab5c-a7f29394c6ac · outbound

This paper cites Basereward: A strong baseline for multimodal reward model.arXiv preprint arXiv:2509.16127, 2025.

Multimodal Reward Hacking in Reinforcement Learning Basereward: A strong baseline for multimodal reward model.arXiv preprint arXiv:2509.16127, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:82afea0413b3eaa709f76e04ae2d676de04c41dec083d2a35809ccd88ab5f633

Observation af9dee90-b69b-43f7-91f5-af013c095807 · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

Multimodal Reward Hacking in Reinforcement Learning MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:7af53b4dafea735c02ca7d1da638c15d6731c239cd89178e305e065b170cfe52

Observation b2e385e7-cee5-4c69-be4b-fe72760b647b · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

Multimodal Reward Hacking in Reinforcement Learning SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:0a6507ef19421ea88686b99465a1464b98dcbb230afeffa90980e811405f750f

Observation 1b5357ef-6ce3-4e4c-ac83-faa07bdb96ad · outbound

This paper cites Generative rlhf-v: Learning principles from multi-modal human preference.Advances in Neural Information Processing Systems, 38: 126021–126051, 2026.

Multimodal Reward Hacking in Reinforcement Learning Generative rlhf-v: Learning principles from multi-modal human preference.Advances in Neural Information Processing Systems, 38: 126021–126051, 2026

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:cfa168148ee7b5d8a195feaba8d1ddcd62cfb41fe4c62f0f60d64a99333d1e44

Observation 9c0eea38-4d5d-4bf1-9421-d4122a1bd893 · outbound

This paper cites Omniguard: Unified safety moderation for omni-modal inputs and outputs.arXiv preprint arXiv:2512.02306, 2025.

Multimodal Reward Hacking in Reinforcement Learning Omniguard: Unified safety moderation for omni-modal inputs and outputs.arXiv preprint arXiv:2512.02306, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:fdd2d4f3a4631ac17c411d35c1472caff3cd53e6ef7965a43ca35c50bb6aae90

Observation 92a31e26-55e3-49f4-9243-7361b6e239b9 · outbound

This paper cites PerPO: Perceptual Preference Optimization via Discriminative Rewarding.

Multimodal Reward Hacking in Reinforcement Learning PerPO: Perceptual Preference Optimization via Discriminative Rewarding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:2afe0679801291ced1a047d816f9dc3074292135aefcdac0e5c5acd6f16b909f

Observation dcb14c7c-e4f6-41f4-9f02-60eedfa91705 · outbound

This paper cites reward hacking.

Multimodal Reward Hacking in Reinforcement Learning reward hacking

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:ce0699f3a7a0ba92afaa1ce53ca822045c5bf190ba88f62f263a715f61f3e8b1

Observation e555e2ac-a863-4f8a-9677-f1caf23279f4 · outbound

This paper cites an unresolved cited work.

Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:17f0d05e3ab92b8ac566ad0abbeeba767f11505c9e03a11b2a350f8a1e599870

Observation b05cf399-1f75-4e76-a7e4-e435aef1a43f · outbound

This paper cites an unresolved cited work.

Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:c751120d687695d5569ef14240f846992224d20b837f49cef091b87ca7207607

Observation c28e247b-dde4-4535-8bf2-02e73727302c · outbound

This paper cites I cannot assist.

Multimodal Reward Hacking in Reinforcement Learning I cannot assist

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:d9ddd36647e915eb8acca2bf4b92db2a70a90225c281ae3e74ed6be834a63709

Observation 064cb242-0dfd-4406-a6f8-84b808095b7a · outbound

This paper cites an unresolved cited work.

Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:7efc522d0d8a9d9f9f463231a34933c97a38c2579ceb298710e54c36432bc13f

Observation 1feeb64e-dc81-4782-9377-6a7085900d28 · outbound

This paper cites label”: “Yes — No — Invalid.

Multimodal Reward Hacking in Reinforcement Learning label”: “Yes — No — Invalid

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:70aeb4e2d2b0ad09f26699fc71af89f52b80a792b67779cc52564305c1966c06

Observation ba1655dd-751d-4878-b080-89785fa014d6 · outbound

This paper cites an unresolved cited work.

Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:eb8103b2a7662b2a106f1c4ba8c11da6e79466efec38b7821441897f8c10eca1

Observation 3c050b4f-3006-4fa4-8d8e-5e6cfef03e44 · outbound

This paper cites an unresolved cited work.

Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:ae92b2f8eafae1e7aff70d8b3d64ecc7b66a257005512fe85edd1c4775fa63f7

Observation 5eaeca36-6819-4e41-a019-bbb477199133 · outbound

This paper cites an unresolved cited work.

Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:c901768625e3a4d335d7fe622cd94a0191e24c5d3358a9dca39b1e0751f47385

Observation e7f5d96d-5fd2-4c19-9a49-a66983734f14 · outbound

This paper cites an unresolved cited work.

Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:54ae70fe152f1426d1745ea7720a1f32d11cf3192d8d38a2e5dea8e0c49d5143

Observation be453a2d-5fed-4e16-9134-5838b9456212 · outbound

This paper cites sft score.

Multimodal Reward Hacking in Reinforcement Learning sft score

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:5ba17c64ae02791879a92136344c8997eb9d45d298a2afc1c5d0739f2a678f6b

Observation aeff1747-e8a9-4de2-b986-59d9dc8b7d03 · outbound

This paper cites an unresolved cited work.

Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:e71f356dd7bcf9e6748c2da472f2480d0d80004868f5a5ffa65266f00d634d11

Observation ddc2ca65-285d-4adf-adf6-73c70c993a70 · outbound

This paper cites I cannot assist. The image contains knives.

Multimodal Reward Hacking in Reinforcement Learning I cannot assist. The image contains knives

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:3d3e9c8fc48f64348430b1498e23c79d85f33b4ba7ef4b867296e2ba80f5fe66

Observation eb47c47a-30e2-4c5d-b2e1-8a47747ea30c · outbound

This paper cites an unresolved cited work.

Multimodal Reward Hacking in Reinforcement Learning Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:cfca0488f0ecabdf2c286e2bf6685f9e42231238eba1992f2f779f76515873ee

Observation 12a10d9a-3a22-4452-8e9f-82be29fc6eff · outbound

This paper cites I cannot help with this request because the image shows instructions for making a weapon, which I cannot assist with.

Multimodal Reward Hacking in Reinforcement Learning I cannot help with this request because the image shows instructions for making a weapon, which I cannot assist with

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:e91a5693fa5dd20eaac9410c4aab46ee6d43d8c6277451a1ba597ebe001aa57a

Pith citing papers

No inbound Pith citation observations are available.