Pith. sign in

Paper Citation Record · LEDGER

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

As of 7 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 13 inbound Pith citation observations for arXiv:2505.17017.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17017 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:56.714854Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:52.913763Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:57.820556Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3e0d43a-e781-4bb8-9f69-4309bee371a1 · outbound

This paper cites https://www.anthropic.com/claude/sonnet/, 2025.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO https://www.anthropic.com/claude/sonnet/, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.913770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:50.317351Z digest=sha256:2e8de3f3b7198554da4cc643b556a4ab3ecb302fe79562cc9b17b8937d0fdfea

Observation 359f5ec5-250e-4b04-a289-1452bb414858 · outbound

This paper cites https://deepmind.google/technologies/gemini/pro/, 2025.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO https://deepmind.google/technologies/gemini/pro/, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.726291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:50.378463Z digest=sha256:332451828ebac0c7e5d4ca38904026716cec56a01ef070f66b0bf022d11a8691

Observation 7f2ec122-7832-494f-8238-9c8fa8bef65c · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.431952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.431952Z digest=sha256:fb0a1589fbd6a4952d6c1fe36f2f6ff30fc116e2be160d84e271c09b16f50c34

Observation 0f70fb72-5c3c-45e7-8989-7a6151e3e5b2 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.485153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.485153Z digest=sha256:dbefbec7c04afa5b3036953b0fabab574d1ee55366dcfb1b276bea9ed955e6a1

Observation 865e8eba-a415-411f-8e38-0a8706c7f078 · outbound

This paper cites Program Synthesis with Large Language Models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Program Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.558588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.558588Z digest=sha256:b8a5ba4ada7bf5662a57206aad57de05a2466e637e91a5ec439804d80356a75d

Observation 9b3f9923-ebb4-4dff-955c-ad8be9d64559 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.610635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.610635Z digest=sha256:a13d4df5861afb4b5d4fa3ffb1281b229b67fc25f70e912e5bad37ba54e4b011

Observation 6b5d2924-ea14-4772-b4dc-cd1c2d13bccb · outbound

This paper cites MaskGIT: Masked generative image transformer.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MaskGIT: Masked generative image transformer

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.581232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:50.698683Z digest=sha256:e39b95786c0997a861757c00a58ebb71df401a3b0f11275d583d5e3f463a6aa2

Observation 73efef0d-ae15-4a8c-ae56-c0b095bb9ebe · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.751073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.751073Z digest=sha256:84d2e7145f695d7970b363bbc8ab1d23d5ef474783b31d2acdb07f460796c185

Observation bf05aa7c-a1ea-4549-8892-0e238d00bbd1 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.825180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.825180Z digest=sha256:5276139e656d3ac8f2066b74a16ef7086442a27add97034de134afb50d87ed01

Observation 9fbdc1ef-c2a2-489d-88c9-3148afa46cd8 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.898025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.898025Z digest=sha256:435b61d5a42db011a0fde4396a3817be67d9dfd6ece61691aec7cd80e2929bc0

Observation ed3149c3-aceb-4027-bb4f-1ecaa2225239 · outbound

This paper cites DeepSeek-R1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DeepSeek-R1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.409929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:50.980864Z digest=sha256:35c714d9d9a7e99110f82e055bc05cd16074d76342a7da5390e5c71fce34ec00

Observation b03dfef3-9e3c-4dc0-85f9-135553a7ca59 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.096255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.096255Z digest=sha256:29765ca1015157f79aebae9db3343e83c220d89f7a64be2c95cb3fc17a6e9537

Observation 507abb59-e777-4a29-a576-b466f32e1ca6 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Taming transformers for high-resolution image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.104750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.104750Z digest=sha256:ab9dfe70119759156938293526c1695abee3c5ab3f699996df0d7f5a61e6dfec

Observation 9cbe2bcc-01ab-4e06-bdce-3af667445360 · outbound

This paper cites Video-R1: Reinforcing video reasoning in mllms, 2025.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Video-R1: Reinforcing video reasoning in mllms, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.209672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:51.147433Z digest=sha256:e1e99e4434a57d93e61effb9d25f8597fabe8d3ade920d584f0b02bea2100d27

Observation c97fdc58-343b-4ef6-adff-f48df04e373a · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.197669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.197669Z digest=sha256:edd03e1b11b67b1eabd65652b6602806a413760830746f4ae816121e97137347

Observation d225a9ee-30c4-4bfb-b251-25ff1b66c382 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.296056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.296056Z digest=sha256:581346723783b9a4d40ea68123e39536a27687b2ad3d9a12252cf68086e5cef4

Observation 217e9aec-6b6f-42ec-b807-9c66c9954767 · outbound

This paper cites SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.415780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.415780Z digest=sha256:0cd04c1715b40736ddf4a052dfab2a931acfa93c2fdbb4c08098826e6ccc680e

Observation 052310a6-09eb-4896-a7dd-72062cfa63af · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.520817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.520817Z digest=sha256:57d2ad6bea7212eedfcba5377e9de134cc70762c1316a7005405e998fd281c98

Observation 2d65c360-5a19-443f-9765-18bcdeb960b2 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Measuring mathematical problem solving with the math dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.647139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.647139Z digest=sha256:a9679effbb793fda9955edd7185943576644eda5cc34a2e27da6b4fc3a0d8219

Observation 245ed011-309c-4e79-b538-9efa1570f074 · outbound

This paper cites Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.020467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:51.759498Z digest=sha256:d01d9d4356afaa0f42bb994f70dbf7718a148f335324d79300d14f6b8a8f923b

Observation 972274fa-eef1-4abd-a1b6-0693320ef687 · outbound

This paper cites T2I-CompBench: A com- prehensive benchmark for open-world compositional text-to-image generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO T2I-CompBench: A com- prehensive benchmark for open-world compositional text-to-image generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:59.888512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:51.859745Z digest=sha256:7c1377242342fd39a6ec87f306aeeeffa602bbaf584ee0e3609fad0b05ea3ba3

Observation d607f266-e310-41ae-8c12-c247a1e5c781 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.953592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.953592Z digest=sha256:df0eceedcc97f2cbb173d1fc5d9d704a436a5c83ffdcad5a6e867d48136dd706

Observation 6cae1717-2a14-4e02-8196-c49abafb20a8 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.031634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.031634Z digest=sha256:d39fdf69e123cefeba73d33f4bda38e43fae01757de5d1aaccc380fbd534144b

Observation 70ee7bc9-52fd-45dd-9250-797e0d1ace69 · outbound

This paper cites Large language models are zero-shot reasoners.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Large language models are zero-shot reasoners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.179275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.179275Z digest=sha256:a9503a4a6935440063a892a403dd9d4673fe45544c08bf7f7e4b221a7af2f8cc

Observation 2be8be33-2dab-4e68-976c-8379b846a7ea · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.344176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.344176Z digest=sha256:196aff3f821f95d3a5f63df8cbe81f46f32bfafcc82926e0882bff15420f5c91

Observation 69914c10-c486-40b0-a882-4fff1238d08c · outbound

This paper cites an unresolved cited work.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.466438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.466438Z digest=sha256:0108571c536071f7f36829959a47966d342ea25bb30a4f1ec7592f4064ecf739

Observation 76765892-f7f3-4f0a-b521-9039ed32dcac · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.582353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.582353Z digest=sha256:75491761a9478e5a10d92043296167b97aff08a379540f23da347e533ccbc1c3

Observation 5c0dda25-23c1-44e5-9915-fb6ca1751178 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.695914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.695914Z digest=sha256:ea8a17611bfb5bb52b09ab615b2fb492ebe314a84950c8a29d07c3df823d97a0

Observation d6931dd5-8f76-4216-ba1f-aaaa4d0fd839 · outbound

This paper cites VideoChat-R1: Enhancing spatio-temporal perception via reinforce- ment fine-tuning, 2025.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoChat-R1: Enhancing spatio-temporal perception via reinforce- ment fine-tuning, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:59.681942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:52.830543Z digest=sha256:5d1dea77f10d12956be1a4e8efb8c82fe7b20f8acb4b0e2662db2fddf787623e

Observation 768dd5fd-8570-4e7a-ad2c-dd089deafb31 · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Cppo: Accelerating the training of group relative policy optimization-based reasoning models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.002220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:53.002220Z digest=sha256:e6bf2ddec9747c1dc9cebea74d02038e7000d744380a145bbe129b72681a0d3b

Observation afd2a74a-c2c3-4465-ba2e-b549029c4eea · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.166352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:53.166352Z digest=sha256:cc652fcd36b094987def9332bacb3a8c38cd055b59152c3d6a0b91bceaf117b2

Observation cd557c76-2f94-4dcb-a8aa-b6415f016d87 · outbound

This paper cites American invitational mathematics examination - aime.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO American invitational mathematics examination - aime

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.258103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:53.258103Z digest=sha256:368eca380e9b7db064561843b86aa6c8d77edc542aa1952d455c2c2ba2416014

Observation 6eddfe43-730f-4c0c-888d-ef2066bd308c · outbound

This paper cites an unresolved cited work.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:59.554579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:53.441712Z digest=sha256:4cd64e45877a59dc381f1df92a83b8dcc75f8574dca8bd74a65ee8d7bde995ec

Observation f8ffe76c-c587-4c75-be09-570240960d69 · outbound

This paper cites Hello gpt-4o.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Hello gpt-4o

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.582237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:53.582237Z digest=sha256:99859c9cf3177bd113e7193191ec4764e95258c90064c91607416131457b5577

Observation c1e073c8-ab00-4c75-822b-d03d10e2ced7 · outbound

This paper cites OpenAI o1 system card, 2024.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO OpenAI o1 system card, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:59.407866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:53.699183Z digest=sha256:252342c6d6797f72b9608c899593f54f2260efd2eaa85addf7b95d716a570e66

Observation 7805cc68-d011-4869-a0ab-e0839c948934 · outbound

This paper cites an unresolved cited work.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.844686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:53.844686Z digest=sha256:5b78129a8b25d9c03412bffe2d9500bb663dac13427460aff768766ad8148f3e

Observation f40e5c07-c350-4479-95d0-67056c5c6e2d · outbound

This paper cites Iterative reasoning preference optimization.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Iterative reasoning preference optimization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:59.219934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:53.982080Z digest=sha256:28838ed31fd9134e14b8dda536f04e5f0437e48fccc0064bf72d749969a3862f

Observation 493ee883-eac6-4594-b315-d8512b81770a · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.154340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.154340Z digest=sha256:d6e577d2eb812bf3c8ab9a05781e6a6e178c91344f610fc5362275827f19f938

Observation 996690e2-b131-4ce5-874d-2a7bc002187f · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Direct preference optimization: Your language model is secretly a reward model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.311678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.311678Z digest=sha256:c9db0acaf5964ebb091be0a5a2374fa1b4edbf5e505b36a8f615bd03c5473541

Observation 8ea8d8dc-3b12-4d22-97fa-e1300cf41597 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO High- resolution image synthesis with latent diffusion models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.426573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.426573Z digest=sha256:eda8756f315fceb54d5b8f1890eb5932aaae0f17199503a29a134d6d6fd447c3

Observation 6246b52d-5954-485e-bd3d-2ebc214a4369 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.498153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.498153Z digest=sha256:50a7c1f06ddbfc7d8d530484dec574eed6d74637e09aacd05fa0de37e7367891

Observation dddb4f89-c7e4-43f6-92ee-f660d2653159 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Proximal Policy Optimization Algorithms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.567116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.567116Z digest=sha256:ca1406b038c55efd758b0921afbf0860dbb32ce4773a566e0763c68cc896f106

Observation 307c3867-8f6f-4974-92fc-8b3b1bc88cd8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.665299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.665299Z digest=sha256:e9e8022d3bb888bbd5571311d179cb5d0396445c4cc74d5d49f00b92a52590f6

Observation 68aa617b-2484-4b7d-8d6d-5518e7f48e6e · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.754970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.754970Z digest=sha256:687f957750b7fa5c13e745a7fa8857b9216fb1fbe14480aaa8523e544e9b13d3

Observation 67ffbb2b-5402-417b-be92-d947d17bcc75 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.845570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.845570Z digest=sha256:a61b4075a10f3ba837e4d831b662b61ca6fbd3c90578daac51868fab0e1fb376

Observation 12979f00-01d9-48cc-b3d3-742632bdd7ac · outbound

This paper cites LaMDA: Language models for dialog applications, 2022.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LaMDA: Language models for dialog applications, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:58.997167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:54.911937Z digest=sha256:0cb7c36cc956d68388679ae629f991011286a95ab873b3cb3ae6ef9686fd5116

Observation 74e00793-1e6b-4c4e-bf61-f5b472092833 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.983985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.983985Z digest=sha256:616416ee170cdb3efd83313c8a4db4c3629dd705e46465de51344786ad41cc9b

Observation 0609fef1-cc26-4865-a8ea-2ac1bf490271 · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.061900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.061900Z digest=sha256:c8fce71369946632485c8d8702bcc50b4e5f6571f53b9f9f61234f20678c3777

Observation 1e187a1a-9e83-4952-b891-dad26567e6e2 · outbound

This paper cites Reasoning in conversation: Solving subjective tasks through dialogue simulation for large language models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Reasoning in conversation: Solving subjective tasks through dialogue simulation for large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:58.794682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:55.131047Z digest=sha256:642331a17bac889a0d7b0a8c96968cf82672b31f1a5c3f708798870ffc436863

Observation 5b9cc19e-30b8-4bc5-b80a-19cec2698a79 · outbound

This paper cites Le, Ed H.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Le, Ed H

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:58.573936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:55.222227Z digest=sha256:a6b3e48ad06c8f92529769c687d048426fad6b3ab681af55c8a8a3b20a1803da

Observation a7820c84-0c1f-4296-9edb-9640123c910d · outbound

This paper cites Unified Reward Model for Multimodal Understanding and Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unified Reward Model for Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.299644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.299644Z digest=sha256:ec8f9045d2800768ee8150de58725281b73c3a82c3df5ea7b95825f9897e298d

Observation 4c0a8dfa-552e-48f1-b390-7b3b6fc9aa2f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Chain-of-thought prompting elicits reasoning in large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.388496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.388496Z digest=sha256:3256f59ef251b00bb61e75615f337bc3494894e63428594ee4cb15d434131e41

Observation 0c95efc2-6a09-4c19-a25d-aecb404deb1b · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.475472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.475472Z digest=sha256:20ce36dd9ad33a79555eb4614ccea296b00ef9d1985015b2d952db0b4c7790b0

Observation 9bbd3866-6666-4c54-bc94-565bacdcece0 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.543870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.543870Z digest=sha256:e894b13f64944c62c1a9cf63d48ce0b23fcde1efc97bd1f668b63989e70efba2

Observation c0e88441-b730-427f-bee8-b976b37589ea · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.602449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.602449Z digest=sha256:ae79399d3a8c0732f535959440c3c3ee733f946db9b250a6c4c05ac6eb71d037

Observation 5f2273de-bae4-489f-9dbd-b4131c80edd4 · outbound

This paper cites ImageReward: Learning and evaluating human preferences for text-to-image generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO ImageReward: Learning and evaluating human preferences for text-to-image generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.664044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.664044Z digest=sha256:6a6660c7a9b5c703f674cbd48142c3033e7a0531f6410bacda70b0035626b490

Observation 2411a786-693a-4729-b382-0ad45550c05d · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.743601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.743601Z digest=sha256:b8c2c32742f959ec34d5f250721379153c6e45fde73d90efe20fad528dd07b25

Observation 111ed8cb-763d-4fa0-8eda-bb83ab36e378 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DanceGRPO: Unleashing GRPO on Visual Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.854961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.854961Z digest=sha256:b4aed5f9dcd3f9ea621cbba6982d1b5685b962b889410a3b88a4e30c42b89d6c

Observation 4ef71de6-d9f6-4576-8362-a8326dc81bc1 · outbound

This paper cites Qwen2 Technical Report.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Qwen2 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.929798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.929798Z digest=sha256:50f7a92f77355c299984843467b0b5d95210b0ceb93c52d2c428787d44089e5e

Observation d5c1b975-fb60-4432-b4ea-9fedb2b99cf2 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Vector-quantized Image Modeling with Improved VQGAN

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.042260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:56.042260Z digest=sha256:a1500676c032fad3a23e1770421bbb1444d21c0884b9b7719ef61297ac588fc7

Observation 249c4d65-1e62-49b1-8920-e501a572bab4 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.192183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:56.192183Z digest=sha256:2da349aae032e00d2d3e8ed6c13c6f78193a5b5a76180b0b617d8d06560bba6c

Observation 88d1fbe9-1623-40cf-bcb3-e30f7da695b7 · outbound

This paper cites ReST-MCTS*: Llm self-training via process reward guided tree search.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO ReST-MCTS*: Llm self-training via process reward guided tree search

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:58.393634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:56.304989Z digest=sha256:c02b0d671c8fa92ae3c4cc1023841af061360515a7ad8d28dbb8f296c6a8cfb6

Observation 32fe3a93-ddcd-4ff6-9d99-efd0f5a22756 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Adding conditional control to text-to-image diffusion models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.411952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:56.411952Z digest=sha256:01247652a88d894d202c4287df96844e0295680990b00ae5906947e1e844f5f6

Observation e1ee7f91-4b22-4b0f-9b66-5ce90750e7d0 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of large language models with zero-initialized attention.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Llama-adapter: Efficient fine-tuning of large language models with zero-initialized attention

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:58.222658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:56.497410Z digest=sha256:5f703d3f34c87cc366aa72e563ffdc3d5c7a68c6569f7b5b114c05abe33128a4

Observation fd2712e2-69dc-46ff-b7c0-86a9b2a7b790 · outbound

This paper cites MathVerse: Does your multi-modal llm truly see the diagrams in visual math problems? ECCV 2024, 2024.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MathVerse: Does your multi-modal llm truly see the diagrams in visual math problems? ECCV 2024, 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:57.995183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:56.612251Z digest=sha256:08cf1bffb0532e1d23a54bbaa755c24f059694d9a8660aa8729dc21aa14f84ae

Observation b529bdf1-551a-4362-aa4e-a75d6d4924ab · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.665148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:56.665148Z digest=sha256:c6b4aedff8935baec533923b4c85274d63350122f3ff51f63961c107afc7e387

Observation feb97369-90bc-4b4f-8c97-12b7ce006be6 · outbound

This paper cites SafetyBench: Evaluating the safety of large language models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SafetyBench: Evaluating the safety of large language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:57.728336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:55:56.714854Z digest=sha256:6f0e3e94970b8875276640f9c9eab3211a7fb0f04d98f2bfe6074832f2d68457

Pith citing papers

Observation 215d5d21-18fc-4a75-a40d-6afad8aef6a6 · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:52.913763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:52.913763Z digest=sha256:7abe98eddb18ac1a6f8c0b5029be988d55210cdbd5cb4c498caeb768253c3483

Observation e5ba5c38-9473-4601-832d-f72cf168f5e3 · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.937303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:79a30330afbd31b7a5d0f286897ea563040d14fa2180e7803d380652b2980bbb

Observation f6720920-3ec0-48c8-ae9e-3d68f885258f · inbound

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE cites this paper.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.062590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:655600def0232e181538059747d47d12b9f199d2cc46a3cd0cc4716cc3e2d7b9

Observation 0751e862-9eb7-42f8-90f8-edd4ff0e57a6 · inbound

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation cites this paper.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.870519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.870519Z digest=sha256:e81eaa780ab43f7e548e11ac986f958350c8edbdff2171701e15a21fafa46d46

Observation 56128ab7-bace-4d99-8095-329180615e9c · inbound

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning cites this paper.

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:31:50.190725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T20:31:28.371368Z digest=sha256:17eab62fd1447ab292e2bfa6a5fbd03b7469e811d438150dc0880c30b04210e5

Observation fa82dde3-e6ba-4cd9-9f27-3430c0e0db5f · inbound

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition cites this paper.

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.397685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:20:58.483350Z digest=sha256:16b457af19c397d15945c377e61e2594379f334dd68350624d16125771571fde

Observation f372d45e-5e7b-40e5-94aa-87ed35887192 · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:04.125091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:04.125091Z digest=sha256:666384320062d8ae4812d66bad455585bff99db105ffbec2b575db6ad715d2d0

Observation d8fe2b4f-c45c-47d2-9105-f294673a21f1 · inbound

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation cites this paper.

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:50:37.149568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T12:50:13.764159Z digest=sha256:24714327ab2e5f1c656d115a4d337f79b9d5ff9e84022a0345d76867c5ae9679

Observation 9fd36647-1579-44e2-9743-98f1d1e369a0 · inbound

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation cites this paper.

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:18:17.195366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T21:14:42.021240Z digest=sha256:59f4035e000bc12053d78145489149b6fa6df0d7044ed61559748d34a8c161f5

Observation ab32f4e7-7889-451a-912b-01d5df6cd009 · inbound

HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation cites this paper.

HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:45:19.190774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T08:42:28.685159Z digest=sha256:7a38e82a181f08b03b54e6d9ed1f642051f7889e590ce2872a4c09d2c3e064c7

Observation d3492cb3-6d89-4e83-9cb4-6f17fc5ff95d · inbound

HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation cites this paper.

HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T22:22:41.696113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T22:22:41.696113Z digest=sha256:64790ec7c893a9ea30614d7e52cffcfda3a9f92f2d61f363e053d3ffe3de2c0b

Observation 5476f687-e9dc-49c4-81ae-c0f378828bc3 · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.824204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:d885b05f7e22eb08aa899afcb8906b08b8452d458bfcc6c93ba97c16a1474494

Observation 7b1e81b3-7ddd-4e79-a312-382928ac04eb · inbound

Optimizing Visual Generative Models via Distribution-wise Rewards cites this paper.

Optimizing Visual Generative Models via Distribution-wise Rewards Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.692094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T16:39:12.711424Z digest=sha256:2a71cd680885c62a16581266710a788115717f25f45ed91c03b8ef583f4ca238