Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:18:06.054981Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2506.03234.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:18:06.054981Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T16:26:34.918099Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T01:37:30.522097Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c232ca51-e49b-4862-ba9d-ffcc20072146 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c5bc4b2-3cf6-451b-8f0e-fa0f87f8f7ff · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Poisoning Attacks against Support Vector Machines
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a7d5036-1e5b-40ac-86eb-e677ced3f55f · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Training Diffusion Models with Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6c3d91-5ef5-4e50-bae9-873226a676fc · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF A survey on generative diffusion models.IEEE Transactions on Knowledge and Data Engineering, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08bfd459-da94-4898-be9a-85e1de397f36 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Trojdiff: Trojan attacks on diffusion models with diverse targets
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 161f05ef-a822-4d0d-9935-86d69a731fc9 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Amplifying membership exposure via data poisoning.Advances in Neural Information Processing Systems, 35:29830– 29844, 2022
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e758bf7b-1171-43c7-83a9-775a9b3d7d14 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Villandiffusion: A unified backdoor attack framework for diffusion models.Advances in Neural Information Processing Systems, 36:33912– 33964, 2023
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d73902d0-311e-4dd5-b6c3-8fe604230f35 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Diffusion models in vision: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9):10850–10869, 2023
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ae54e1-c492-41d2-9d81-98f72add11b5 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF A survey on data poisoning attacks and defenses
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 154ec68f-8638-40af-9e34-1f2dc9129a4a · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Dpok: Reinforcement learning for fine-tuning text-to-image diffusion models.Advances in Neural Information Processing Systems, 36:79858–79885, 2023
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d7a39fa-b174-45d3-9f5e-7a4c4ddc7b8f · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36:36652–36663, 2023
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b37bad84-ff1b-432d-a2b4-af5336bffd74 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Aligning Text-to-Image Models using Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae670811-d078-4079-a6b9-4f6066a54382 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f08532-d743-47ae-a892-36c3f5f6222f · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2d3d101-2de6-41d1-ba22-f1a52ee1920b · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Backdooring Bias ($B^2$) into Stable Diffusion Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f949655-be9a-4032-8c9f-09a7e5a7ad61 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF From trojan horses to castle walls: Unveiling bilateral data poisoning effects in diffusion models.Advances in Neural Information Processing Systems, 37:82265–82295, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59a7cbec-dd05-4dad-bc13-825bbfec08d6 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Learning transferable visual models from natural language supervision
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6007439-ac16-4601-bf79-aefa8c707234 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Universal Jailbreak Backdoors from Poisoned Human Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50d6e421-ac7a-49fb-9868-0ca8442a5a76 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e05d140-3491-48fb-814f-a6e74ae761a1 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Poison frogs! targeted clean-label poisoning attacks on neural networks.Advances in neural information processing systems, 31, 2018
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54f7eae3-a740-4c7a-ad9f-5aac9f7f5c47 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Nightshade: Prompt-specific poisoning attacks on text-to-image generative models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9380d697-33e0-47d7-8aa8-f39f48fe80e4 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Defining and characterizing reward gaming.Advances in Neural Information Processing Systems, 35:9460– 9471, 2022
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff8d6a03-ecc2-4516-9809-9da4766a8a65 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Attacks and defenses for generative diffusion models: A comprehensive survey.ACM Computing Surveys, 57(8):1–44, 2025
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c35545e0-c05a-4570-9d41-3a9605c58095 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Rlhfpoi- son: Reward poisoning attack for reinforcement learning with human feedback in large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7e337ca-da47-4492-90f3-2a6be172ac8c · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Preference Poisoning Attacks on Reward Model Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c07524e-4b2d-4435-b510-47ce1c7e16e8 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5955355f-aeb3-4f18-a634-ce49693845f4 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Human preference score: Better aligning text-to-image models with human preference
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 180fd271-a4cc-4700-9b6f-c017fbcedf9d · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Adversarial label flips attack on support vector machines
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd29907a-9f14-49c8-b5f2-e187b32ce0d1 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Imagereward: Learning and evaluating human preferences for text-to-image generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 859bd800-3a36-4ae6-9149-195f24f1cdce · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e6ebbea-9cba-46fe-a5b7-8c55c5a571cc · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Using human feedback to fine-tune diffusion models without any reward model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65970768-5df4-4d28-8b38-2231aaf7b64e · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Diffusion models: A comprehensive survey of methods and applications.ACM Computing Surveys, 56(4):1–39, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3442542c-30c4-4352-af91-c631d8821e31 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Poisonprompt: Backdoor attack on prompt-based large language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0341263c-6b8a-4661-8ce4-95d8a6a8b82d · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Text- to-image diffusion models can be easily backdoored through multimodal data poisoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e926ed-4c2b-437d-bc93-608708a72727 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Text-to-image Diffusion Models in Generative AI: A Survey
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b07a20c-d990-48b9-a212-dd7b6110fd25 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Aligning few-step diffusion models with dense reward difference learning.arXiv preprint arXiv:2411.11727, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e09bfbe4-338b-43b7-8659-9d26de7da513 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Shielding collaborative learning: Mitigating poisoning attacks through client-side detection.IEEE Transactions on Dependable and Secure Computing, 18(5):2029–2041, 2020
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9268e5d-894a-41da-860e-3e09c4e33196 · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Diffusion Models for Reinforcement Learning: A Survey
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c89c4aad-343c-4484-8241-cfbb1765f46e · outbound
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b6439e1-8155-4c28-8686-855b5bee67a9 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
Reference 221
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b58c985f-04cf-476f-ab99-e55e1c7a5c69 · inbound
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2607aec9-0022-4ee2-95b0-b7697c054fc0 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
Reference 295
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.