Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:56.714854Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 13 inbound Pith citation observations for arXiv:2505.17017.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:56.714854Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:52.913763Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:18:57.820556Z
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c3e0d43a-e781-4bb8-9f69-4309bee371a1 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO https://www.anthropic.com/claude/sonnet/, 2025
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 359f5ec5-250e-4b04-a289-1452bb414858 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO https://deepmind.google/technologies/gemini/pro/, 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f2ec122-7832-494f-8238-9c8fa8bef65c · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f70fb72-5c3c-45e7-8989-7a6151e3e5b2 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 865e8eba-a415-411f-8e38-0a8706c7f078 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Program Synthesis with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b3f9923-ebb4-4dff-955c-ad8be9d64559 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b5d2924-ea14-4772-b4dc-cd1c2d13bccb · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MaskGIT: Masked generative image transformer
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73efef0d-ae15-4a8c-ae56-c0b095bb9ebe · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Evaluating Large Language Models Trained on Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf05aa7c-a1ea-4549-8892-0e238d00bbd1 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fbdc1ef-c2a2-489d-88c9-3148afa46cd8 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed3149c3-aceb-4027-bb4f-1ecaa2225239 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DeepSeek-R1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b03dfef3-9e3c-4dc0-85f9-135553a7ca59 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Scaling rectified flow trans- formers for high-resolution image synthesis
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 507abb59-e777-4a29-a576-b466f32e1ca6 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Taming transformers for high-resolution image synthesis
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cbe2bcc-01ab-4e06-bdce-3af667445360 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Video-R1: Reinforcing video reasoning in mllms, 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c97fdc58-343b-4ef6-adff-f48df04e373a · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Geneval: An object-focused framework for evaluating text-to-image alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d225a9ee-30c4-4bfb-b251-25ff1b66c382 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217e9aec-6b6f-42ec-b807-9c66c9954767 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 052310a6-09eb-4896-a7dd-72062cfa63af · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d65c360-5a19-443f-9765-18bcdeb960b2 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Measuring mathematical problem solving with the math dataset
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245ed011-309c-4e79-b538-9efa1570f074 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 972274fa-eef1-4abd-a1b6-0693320ef687 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO T2I-CompBench: A com- prehensive benchmark for open-world compositional text-to-image generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d607f266-e310-41ae-8c12-c247a1e5c781 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cae1717-2a14-4e02-8196-c49abafb20a8 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ee7bc9-52fd-45dd-9250-797e0d1ace69 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Large language models are zero-shot reasoners
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2be8be33-2dab-4e68-976c-8379b846a7ea · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69914c10-c486-40b0-a882-4fff1238d08c · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76765892-f7f3-4f0a-b521-9039ed32dcac · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LLaVA-OneVision: Easy Visual Task Transfer
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c0dda25-23c1-44e5-9915-fb6ca1751178 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6931dd5-8f76-4216-ba1f-aaaa4d0fd839 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoChat-R1: Enhancing spatio-temporal perception via reinforce- ment fine-tuning, 2025
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 768dd5fd-8570-4e7a-ad2c-dd089deafb31 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Cppo: Accelerating the training of group relative policy optimization-based reasoning models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afd2a74a-c2c3-4465-ba2e-b549029c4eea · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd557c76-2f94-4dcb-a8aa-b6415f016d87 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO American invitational mathematics examination - aime
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eddfe43-730f-4c0c-888d-ef2066bd308c · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8ffe76c-c587-4c75-be09-570240960d69 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Hello gpt-4o
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e073c8-ab00-4c75-822b-d03d10e2ced7 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO OpenAI o1 system card, 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7805cc68-d011-4869-a0ab-e0839c948934 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f40e5c07-c350-4479-95d0-67056c5c6e2d · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Iterative reasoning preference optimization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 493ee883-eac6-4594-b315-d8512b81770a · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 996690e2-b131-4ce5-874d-2a7bc002187f · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Direct preference optimization: Your language model is secretly a reward model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea8d8dc-3b12-4d22-97fa-e1300cf41597 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO High- resolution image synthesis with latent diffusion models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6246b52d-5954-485e-bd3d-2ebc214a4369 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dddb4f89-c7e4-43f6-92ee-f660d2653159 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Proximal Policy Optimization Algorithms
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 307c3867-8f6f-4974-92fc-8b3b1bc88cd8 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68aa617b-2484-4b7d-8d6d-5518e7f48e6e · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ffbb2b-5402-417b-be92-d947d17bcc75 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12979f00-01d9-48cc-b3d3-742632bdd7ac · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LaMDA: Language models for dialog applications, 2022
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74e00793-1e6b-4c4e-bf61-f5b472092833 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LLaMA: Open and Efficient Foundation Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0609fef1-cc26-4865-a8ea-2ac1bf490271 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e187a1a-9e83-4952-b891-dad26567e6e2 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Reasoning in conversation: Solving subjective tasks through dialogue simulation for large language models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b9cc19e-30b8-4bc5-b80a-19cec2698a79 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Le, Ed H
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7820c84-0c1f-4296-9edb-9640123c910d · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unified Reward Model for Multimodal Understanding and Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c0a8dfa-552e-48f1-b390-7b3b6fc9aa2f · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Chain-of-thought prompting elicits reasoning in large language models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c95efc2-6a09-4c19-a25d-aecb404deb1b · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bbd3866-6666-4c54-bc94-565bacdcece0 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0e88441-b730-427f-bee8-b976b37589ea · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f2273de-bae4-489f-9dbd-b4131c80edd4 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO ImageReward: Learning and evaluating human preferences for text-to-image generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2411a786-693a-4729-b382-0ad45550c05d · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111ed8cb-763d-4fa0-8eda-bb83ab36e378 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DanceGRPO: Unleashing GRPO on Visual Generation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef71de6-d9f6-4576-8362-a8326dc81bc1 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Qwen2 Technical Report
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5c1b975-fb60-4432-b4ea-9fedb2b99cf2 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Vector-quantized Image Modeling with Improved VQGAN
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 249c4d65-1e62-49b1-8920-e501a572bab4 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d1fbe9-1623-40cf-bcb3-e30f7da695b7 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO ReST-MCTS*: Llm self-training via process reward guided tree search
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32fe3a93-ddcd-4ff6-9d99-efd0f5a22756 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Adding conditional control to text-to-image diffusion models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1ee7f91-4b22-4b0f-9b66-5ce90750e7d0 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Llama-adapter: Efficient fine-tuning of large language models with zero-initialized attention
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd2712e2-69dc-46ff-b7c0-86a9b2a7b790 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MathVerse: Does your multi-modal llm truly see the diagrams in visual math problems? ECCV 2024, 2024
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b529bdf1-551a-4362-aa4e-a75d6d4924ab · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb97369-90bc-4b4f-8c97-12b7ce006be6 · outbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SafetyBench: Evaluating the safety of large language models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 215d5d21-18fc-4a75-a40d-6afad8aef6a6 · inbound
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ba5c38-9473-4601-832d-f72cf168f5e3 · inbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6720920-3ec0-48c8-ae9e-3d68f885258f · inbound
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0751e862-9eb7-42f8-90f8-edd4ff0e57a6 · inbound
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56128ab7-bace-4d99-8095-329180615e9c · inbound
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa82dde3-e6ba-4cd9-9f27-3430c0e0db5f · inbound
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f372d45e-5e7b-40e5-94aa-87ed35887192 · inbound
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8fe2b4f-c45c-47d2-9105-f294673a21f1 · inbound
From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9fd36647-1579-44e2-9743-98f1d1e369a0 · inbound
VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab32f4e7-7889-451a-912b-01d5df6cd009 · inbound
HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3492cb3-6d89-4e83-9cb4-6f17fc5ff95d · inbound
HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5476f687-e9dc-49c4-81ae-c0f378828bc3 · inbound
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b1e81b3-7ddd-4e79-a312-382928ac04eb · inbound
Optimizing Visual Generative Models via Distribution-wise Rewards Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.