Pith. sign in

Paper Citation Record · LEDGER

Towards Self-Improvement of Diffusion Models via Group Preference Optimization

As of 18 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 3 inbound Pith citation observations for arXiv:2505.11070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11070 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:03:31.022584Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T01:22:42.691009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T01:23:26.825361Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e953848-7065-4f72-ada8-c6169d4c3a71 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.102049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.102049Z digest=sha256:9395ccb84fa898db7fe296c3aefec935aa2a85d8f8b283c6888a34401b5c8231

Observation 7847d3a9-e927-45de-a57d-1f04bb30df03 · outbound

This paper cites A Noise is Worth Diffusion Guidance.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization A Noise is Worth Diffusion Guidance

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.108005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.108005Z digest=sha256:d0b26e0576fcdb93b02c2cf9828bc15551ac1ea4be82c9dd7bd1cf9eb438e652

Observation 7b758ff8-c5b1-43df-a53c-5c8df4595df2 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.170314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.170314Z digest=sha256:848ba40ed57335cf0fccd5cd645a4ae68fbc69dde5d45748bfbe9a47a5a36957

Observation 274f9632-d4b2-4cbe-97ea-3eca80d4da9d · outbound

This paper cites Improving image generation with better captions.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.316991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.316991Z digest=sha256:9d05fc81274b3cb9ea761f85987c24f429884ebdc2ffbd336d5fd30c76f5d4bf

Observation 13f56677-9437-4626-9d64-409c9deda726 · outbound

This paper cites Make It Count: Text-to-Image Generation with an Accurate Number of Objects.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Make It Count: Text-to-Image Generation with an Accurate Number of Objects

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.329115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.329115Z digest=sha256:101d9124bf1d2331bf3c6fbae59aedc5db7a4a00b22d260b6f7668cc1421e61f

Observation b09b675c-fb17-4f90-aa27-2c290cb92c9b · outbound

This paper cites Training diffusion models with reinforcement learning.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Training diffusion models with reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:33.995342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:29.334450Z digest=sha256:e473991fad0b36dcc6e5d96b174e66bf3f508e05535436fb9bc3cf48d64cf46f

Observation ee70a06d-53a8-4807-8b65-3af4aff2b3cd · outbound

This paper cites Text-to-image diffusion models cannot count, and prompt refinement cannot help.arXiv preprint arXiv:2503.06884, 2025.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Text-to-image diffusion models cannot count, and prompt refinement cannot help.arXiv preprint arXiv:2503.06884, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.476120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.476120Z digest=sha256:b628e6056a9776e55f46eca2da734e17a0c2e66063c858f7e1eef6c840a7f27f

Observation 9cebc991-728a-4b46-9925-428d474a7f4d · outbound

This paper cites Getting it right: Improving spatial consistency in text-to-image models.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Getting it right: Improving spatial consistency in text-to-image models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.523559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.523559Z digest=sha256:c2b49ee084b98caba70880c7f95137060096131fe163380dab77dfae35f3795f

Observation 85053dad-5263-49cc-a503-64ffad5c5e10 · outbound

This paper cites Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36: 9353–9387, 2023.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36: 9353–9387, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.527769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.527769Z digest=sha256:2b748472dc5436c9b35a9ad0d0b2ee2acb4f7d6074f7aea38ed7448900815526

Observation 490fc261-c788-4740-81df-2631e3e86dca · outbound

This paper cites Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.532979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.532979Z digest=sha256:9a0051ea44b808566284757005ba97d16283e81c4b2beb19882ee0ea9a500906

Observation 7974a9e4-7d50-4056-b250-d96f0e658b5b · outbound

This paper cites Directly fine-tuning diffusion models on differentiable rewards.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Directly fine-tuning diffusion models on differentiable rewards

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:33.828398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:29.564250Z digest=sha256:79b8c6f5d0ff9a42998464bfdbf3689429b93002e168b4a2b197bedae1c87ea9

Observation fd9813fa-7e29-4c47-b405-ff5940ae63aa · outbound

This paper cites Less is more: Improving llm alignment via preference data selection.arXiv preprint arXiv:2502.14560, 2025.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Less is more: Improving llm alignment via preference data selection.arXiv preprint arXiv:2502.14560, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.657047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.657047Z digest=sha256:ebd01309006966077596b005df4f7511ea4bbe46566b79b70f57966b8f60b987

Observation 802229b2-2d11-4950-90a3-c2a389ac29fd · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.682007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.682007Z digest=sha256:1df07af75398bd650edafd7df63f1065366891a3b2360665cd28c6dad5fc578f

Observation 4e94ab5b-e4bd-4edb-be33-6c4e400249d9 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization KTO: Model Alignment as Prospect Theoretic Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.687518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.687518Z digest=sha256:dd4e6178f23c08e84c285a9737812a0d32fca65e4961bed724e16c0ab415b656

Observation 9cf19877-b93a-433e-9dc7-06259cc06206 · outbound

This paper cites Denoising diffusion probabilistic models.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Denoising diffusion probabilistic models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:33.804017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:29.692678Z digest=sha256:ce9d482db80d3a52357b61ddb7381178db619b4ffa56a427197a4e1c43201240

Observation 3933286f-f058-40f8-a72c-4d45250068da · outbound

This paper cites Reference-free monolithic preference optimization with odds ratio.arXiv e-prints, pages arXiv–2403, 2024.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Reference-free monolithic preference optimization with odds ratio.arXiv e-prints, pages arXiv–2403, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:33.697525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:29.698201Z digest=sha256:9bea200adc28441ee4392184911482c92b7bb94841d980e54cec6569d45ab671

Observation 00fd11c1-05f9-4d37-b7ef-a67ee53eacd0 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.788600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.788600Z digest=sha256:c2210b2f658a3ea23f72b9c2ca228397920ad25ce239030e83816562e7a3a06f

Observation d6648fcc-ffa9-4c7e-8497-f112840fc5cb · outbound

This paper cites T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation.Advances in Neural Information Processing Systems, 36:78723–78747, 2023.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation.Advances in Neural Information Processing Systems, 36:78723–78747, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.831361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.831361Z digest=sha256:8ea090e3c04d20fc77d3f722fe52e5679f144aa15c3d61883a045dd46a99b771

Observation 99bf66c8-fe77-45bc-b392-8f90721ccfea · outbound

This paper cites Ultralytics YOLO, 2023.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Ultralytics YOLO, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:33.673547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:29.836368Z digest=sha256:385b7a5e317f79d51dd7c3a5e81d9c1423a09a0685af905b789d2846076794f8

Observation fa1324b2-863e-477f-ab83-6a35503a4b45 · outbound

This paper cites Scalable Ranked Preference Optimization for Text-to-Image Generation.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Scalable Ranked Preference Optimization for Text-to-Image Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.841117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.841117Z digest=sha256:245885ca38fb2384292085331f21366900d24527e3d6ca88cfd8ac06c0e920ae

Observation 41ef1a48-bd0c-432d-9b17-adaa98cef521 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36:36652–36663, 2023.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36:36652–36663, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.845960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.845960Z digest=sha256:ee23dda7b994f3d48f622de31fda5d6101accb48dd2a4afa431786873af72e20

Observation 753bfbe8-e6be-47c9-8903-f0ad461da02b · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2024.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Flux.https://github.com/black-forest-labs/flux, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.869669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.869669Z digest=sha256:34b5c4284f7b0de826878ff89168c8b40f035c5401f84b294fae88d396fc00e6

Observation 9a828dea-16ed-4a1a-848b-d20c5775bb74 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Aligning Text-to-Image Models using Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:29.937680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:29.937680Z digest=sha256:de45674c6f1ae48f07c591212e19e990ca127425ee995d76aaf8ff3c98c93631

Observation 51015cd7-63b9-4238-8436-75b259f13879 · outbound

This paper cites Calibrated multi-preference optimization for aligning diffusion models.arXiv preprint arXiv:2502.02588, 2025.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Calibrated multi-preference optimization for aligning diffusion models.arXiv preprint arXiv:2502.02588, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.002767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.002767Z digest=sha256:a44590b51ea3060c6ec56e2a30bece4e2037ebb008f15afefe321a8919a8d00f

Observation f4ee8f8d-cac5-4bfd-a506-c99538d45ba6 · outbound

This paper cites Align- ing diffusion models by optimizing human utility.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Align- ing diffusion models by optimizing human utility

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:33.338641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.008103Z digest=sha256:b7f02f6808d4ca70972379033de141c454d7a5e410f85c4adbbabadddc045bcd

Observation 692af499-7b68-439b-b86f-5f343ccd17bd · outbound

This paper cites Policy Optimization in RLHF: The Impact of Out-of-preference Data.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Policy Optimization in RLHF: The Impact of Out-of-preference Data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.012249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.012249Z digest=sha256:234e3e9ce9f352a8459670799ee81c6280e239a035504411210ce96f899563f5

Observation eb14da43-0d95-488c-9015-fa9e59c6b20a · outbound

This paper cites Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.017303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.017303Z digest=sha256:530ad57d00c44a423401bb02df862c9d0a64a4413a0ca5951d7ddd52427a7a92

Observation 129297b4-09fd-40ea-a54c-ae81a42d30ca · outbound

This paper cites Flow matching for generative modeling.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Flow matching for generative modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.106144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.106144Z digest=sha256:6c64512b6c9cff20ac3a52d6db79d4fd5d79fcd6e3aa1113a6cbd659818c9450

Observation 347b89c1-dce8-4a8c-b6bd-4879a2c168b5 · outbound

This paper cites Alignment of diffusion models: Fundamentals, challenges, and future.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Alignment of diffusion models: Fundamentals, challenges, and future

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.142072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.142072Z digest=sha256:ef37b41af8e248c9319ca1dca36f7dd5cbba8e0b6e6c7f8cd973b4d9dce4d341

Observation ad9a4394-030f-4035-8356-659cde4c78f7 · outbound

This paper cites Character-Aware Models Improve Visual Text Rendering.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Character-Aware Models Improve Visual Text Rendering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.147094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.147094Z digest=sha256:3406760aaa77dc41a42f036714ca0b0218d49281b4c27f2262bba1018e3dfec6

Observation 7355a2d9-b7e8-42e2-91fd-ff64e0913867 · outbound

This paper cites VideoDPO: Omni-Preference Alignment for Video Diffusion Generation.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization VideoDPO: Omni-Preference Alignment for Video Diffusion Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.152272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.152272Z digest=sha256:8fd9f0d4c97343981dde68ecb134b32589a8016345c94fffef1bcf76c9d14f70

Observation 2f1472b5-1879-433b-a93f-c5ea1d7bd393 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.191853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.191853Z digest=sha256:7f71089bf900d8a8da15839ef3a690dad7cb1cb161b405c5b0623cd238a5430e

Observation 69c6ac3c-ca9d-42eb-b226-cdc85e0e8b1b · outbound

This paper cites Decoupled Weight Decay Regularization.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Decoupled Weight Decay Regularization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.282842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.282842Z digest=sha256:ed8cbe10bad6df5fffd4a451797fcf2ff32db10c221507c127dea4b945f664d8

Observation 48a563b7-d952-4734-8505-5712db092a4c · outbound

This paper cites Exploring the role of large language models in prompt encoding for diffusion models.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Exploring the role of large language models in prompt encoding for diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:33.179537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.307954Z digest=sha256:b73d3be6ac259702ffbac1f4db463793eb2bfc16db134ad396aa798fff02df2a

Observation 17db391e-f2af-416f-b3ca-a10e8e8a9c16 · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.313409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.313409Z digest=sha256:3e0032f863ee06eaec3d46364bdd654c094e97b88db62830b63cabb86363933e

Observation 6089de64-5d4e-41ce-83e3-739a35e9b911 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.318162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.318162Z digest=sha256:3c8ec09c550b092bdde945784c8f23d6dabf4c766c8e5857e104f310fd0666ce

Observation 5058c826-aea0-4f94-9d09-1cee8912daf0 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.322997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.322997Z digest=sha256:121fd121e516e3adaf1852a5f27525b8352b98d852309d757df5a31263e5a6f7

Observation fd28208b-e9c4-42b2-b0aa-2e5254280c9e · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization High- resolution image synthesis with latent diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.366530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.366530Z digest=sha256:12fe454aec4ce1661ff753eb7c65cc27e07e4998ff02a4df53482487ec0e2f97

Observation 8a72b7d1-e6df-4507-a3cf-eeb20fc67e37 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.446989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.446989Z digest=sha256:71254281a54772a7f2412e584d4acd401eaf490ee681ceb88e2fb3081d995797

Observation 45d2b281-0a38-48fa-99fa-5ccfd26354f2 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.485008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.485008Z digest=sha256:0e6eacdf0e3b78c0078aca746017529c32a9431a4e3a03fe8e34df8e6776670d

Observation e6823a61-d656-4790-8d6c-9655808ba213 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.489775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.489775Z digest=sha256:830f53cc186c7fa7d4ec6572c4034a12c6cd075e2766c676efba07ab95c613c6

Observation d89ee5c9-9d03-4f0d-90ac-c25e60b3c2d4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.494204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.494204Z digest=sha256:c49c2136c957e8030da8e3dcf650415b7b471f2e1d43f5e594c55c9c72b0bb18

Observation 2e736897-9f34-4cb9-bfd4-96b931107ada · outbound

This paper cites Anytext: Multilingual visual text generation and editing.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Anytext: Multilingual visual text generation and editing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:33.022200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.498711Z digest=sha256:4124ea3f63ca4ecfe45db4c617725897c2b6b1a7d03a720a72a4ecbdb7fcd362

Observation 3a6b9c14-9c8b-455d-b7f5-01fb6c82dd8c · outbound

This paper cites Diffusion model alignment using direct preference optimization.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Diffusion model alignment using direct preference optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.503325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.503325Z digest=sha256:8c74d4549d29e59991717f147db1bc16acdd98995486ac1a72e95c42a755704a

Observation 6cb75f59-a21e-4fa2-9456-d42ff7802ef0 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Wan: Open and Advanced Large-Scale Video Generative Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.507321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.507321Z digest=sha256:89730cfcb7ff95896362a0ca87290e1feb5a6bf0fb1aedc2d74fbe3adc196dd0

Observation abf9200d-ece3-44ff-b486-f4709ca2e5f3 · outbound

This paper cites The Silent Assistant: NoiseQuery as Implicit Guidance for Goal-Driven Image Generation.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization The Silent Assistant: NoiseQuery as Implicit Guidance for Goal-Driven Image Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.511812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.511812Z digest=sha256:506c28154ec7db0503f42889de9c35beb82e07f6daad1eea7d41c7283ca0c998

Observation 02f14d9f-2060-4f2c-8819-a6aa77223df2 · outbound

This paper cites DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.516736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.516736Z digest=sha256:669db354207e211fa54f8305fc1f79d5bb77de81609ad14d8e0c8a7db3d1ca7f

Observation 3664726c-4fec-46f3-837d-d7a6c116194f · outbound

This paper cites Human preference score: Better aligning text-to-image models with human preference.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Human preference score: Better aligning text-to-image models with human preference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.567176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.567176Z digest=sha256:25cdb6a496997f509e136421b19c20fabf4a8e00d2a49f73f90833b444a0763f

Observation 369312ab-7a59-462a-8f81-bc13e6d06e98 · outbound

This paper cites Deep reward supervisions for tuning text-to-image diffusion models.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Deep reward supervisions for tuning text-to-image diffusion models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:32.870282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.626293Z digest=sha256:29f219bd911eaf139032d0235bce8d676ef8eb309112afdf77d5640c57d16421

Observation 7f38f5e4-7a63-41a5-bca5-9f0599f5fecd · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Imagereward: Learning and evaluating human preferences for text-to-image generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.722215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.722215Z digest=sha256:976f6233408cfa5e36b102691ff71275a72d7e34c8e449deddd8b294b43f9dac

Observation 21e201c2-65be-41fe-8724-13950d14673c · outbound

This paper cites Using human feedback to fine-tune diffusion models without any reward model.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Using human feedback to fine-tune diffusion models without any reward model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.772282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.772282Z digest=sha256:102365dc7f371f1464f55ba71c3c18b285c65295290aab8042027b002b85bd01

Observation 10b0b9e1-be8f-45db-ae76-2ea878cc25c6 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Adding conditional control to text-to-image diffusion models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.811452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.811452Z digest=sha256:c0dce8608ac99144e216bef5a608994e9375a42c5cb32fa75535c26741b9e489

Observation a7861898-c017-4067-bb12-1a43efed4515 · outbound

This paper cites Learning multi-dimensional human preference for text-to-image generation.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Learning multi-dimensional human preference for text-to-image generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.816995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.816995Z digest=sha256:a1e0d621a9e4dcf1b7172a07f341c0bac86acb71dd606291a84437d09e78dc0c

Observation 69be1f60-03dd-406b-b8ec-019c665d638d · outbound

This paper cites Diffusion model as a noise-aware latent reward model for step-level preference optimization.arXiv preprint arXiv:2502.01051, 2025.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Diffusion model as a noise-aware latent reward model for step-level preference optimization.arXiv preprint arXiv:2502.01051, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.820896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.820896Z digest=sha256:8dc318555043f2973850026e9e256cd64b93303c29129b0e2c0bf8bd1f06815c

Observation dc874733-eb87-47e4-b933-973e3c9ca48f · outbound

This paper cites Cogview3: Finer and faster text-to-image generation via relay diffusion.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Cogview3: Finer and faster text-to-image generation via relay diffusion

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:32.752256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.825317Z digest=sha256:0a441c7c8f6794368cfcc9d4f660d97e4253fbd4cb49dc40263a2bb4018d7b7a

Observation 95c13eee-ed9d-4de1-a004-14e332433c13 · outbound

This paper cites Golden Noise for Diffusion Models: A Learning Framework.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Golden Noise for Diffusion Models: A Learning Framework

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:30.829602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:30.829602Z digest=sha256:9abfa254e8bb38fc57b184736f24db0701477b6f1d2bb1cbcad24b44d9cb5ab1

Observation 7c98595b-93a2-4b0e-9e5b-33e3db25fbe5 · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.548329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.834830Z digest=sha256:f1f713974a2320c04645a65e2748457ea8b05947741990381965d5854141a39e

Observation 64ab745c-29ef-4a14-85cb-0f4cd9dfd25f · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.534768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.838992Z digest=sha256:d6fb6e1dbed406a3949d3902ac2ea7a6ae9231531f647e5d3f4faa5244b2c624

Observation fe7e4b3a-0210-4be2-97b5-eca49a76c89d · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.389993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.843453Z digest=sha256:741e32f4a0aa695d60db288e4accbd3a2e7c276b80e6f58447d606154b2bc4db

Observation d07bf9d1-835b-4e23-97b1-aa17533e6cad · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.376602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.847851Z digest=sha256:bfe91a97facc34aac5de10028e71ca5b014c33778afc3a3270f609e596da4013

Observation 1433c578-539f-46ad-b25b-ae04574bf9f6 · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.289980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.857626Z digest=sha256:0767155c00641c82367b5f306d197905247842e3e6391b19b24382a94668f652

Observation b771aa8f-e089-4911-85f6-55c2d7047938 · outbound

This paper cites [Example Template] Input: 3 cat Output: Three cats curled up together on a sunny windowsill.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization [Example Template] Input: 3 cat Output: Three cats curled up together on a sunny windowsill

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:32.146541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.862829Z digest=sha256:9e02e356173d1043579743d050090396e9f77535f0ce864a4fa12030fb9c0ee2

Observation cfc0a1e3-6450-4c63-af16-b93324284294 · outbound

This paper cites Follow these guidelines: [Output Requirements] Generate prompts with this structure:.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Follow these guidelines: [Output Requirements] Generate prompts with this structure:

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:32.070189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.988425Z digest=sha256:169bf4bab419f6105773bcb9623bf069ab558f84797cd7d183ca1e09537f43e5

Observation 8795ca4e-7c45-45c4-ba0f-35930758bef0 · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.054519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.992987Z digest=sha256:092bb1c182fc91b0e64a2e526509f17ec5a501a6c2b51ed5f64698997dce82b3

Observation 03d3efa6-e140-40c1-824d-4639cd532fb1 · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.362009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:30.997488Z digest=sha256:91a763d08fabac7bed58273dc7a434a1666094b4fb4d74332c12f9eb044ee12a

Observation 546442cd-8ab0-4243-b532-2b14d0d3ccef · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.038270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:31.001822Z digest=sha256:f76ace8ca3a52334b456afbbc6b7ff53618edc5af6c440ea3774f687f4051831

Observation ed76670d-67ac-44f9-adf9-a9a1398d391c · outbound

This paper cites [Optimization Principles].

Towards Self-Improvement of Diffusion Models via Group Preference Optimization [Optimization Principles]

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:31.810325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:31.006032Z digest=sha256:56d2c2be15280104176d80cd013ee482d64d765851ee72788263e72935404f82

Observation 3f7af134-3d70-4f97-a6fe-0c11e796209c · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.114769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:31.009809Z digest=sha256:8ff5b31b0b21c235d70159bbba0eaa14029776ff2db386bdfa667a37b63819ce

Observation e389d7ff-09e1-44da-8ef1-406603d085b6 · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.100930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:31.013913Z digest=sha256:3999809ec9a0be55e078cc2f07492cd0983c3268a2eff39da3b7c4f42c9acc19

Observation 0606f03b-ce42-4e69-8271-8dae054c9e40 · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:32.085862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:31.018464Z digest=sha256:20aa8dd308eaae2b5a32319d645ceb3ecbf8e8160e5f1c450ccf6f9219e807a7

Observation 025ed2ad-7697-43b0-b4d3-de71033cffcd · outbound

This paper cites an unresolved cited work.

Towards Self-Improvement of Diffusion Models via Group Preference Optimization Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:31.771090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:31.022584Z digest=sha256:3159911eda2ae1e263dd741d102eb3006e20f6a7c47b60418c07595151da1139

Pith citing papers

Observation 6cbf8414-1296-427e-8e31-d67b50a322cd · inbound

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal cites this paper.

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal Towards Self-Improvement of Diffusion Models via Group Preference Optimization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:26.827986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T01:22:42.691009Z digest=sha256:b3bd4ecc32e0b0acc033fb5197d307371ea5913320b0f52767e2d6dd513acd2d

Observation d933ab5c-a43a-49a1-a136-38b95e3dbd6c · inbound

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling cites this paper.

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling Towards Self-Improvement of Diffusion Models via Group Preference Optimization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:07.642353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:10:06.994557Z digest=sha256:f30ccb641e026d2e4c00c620710d29973d39b3f415c7f1e4e38ddbd5628ff7f7

Observation 3cf3dc95-2919-4c21-b3b7-80f0b4092774 · inbound

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories cites this paper.

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories Towards Self-Improvement of Diffusion Models via Group Preference Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:25:18.862887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T11:23:28.424453Z digest=sha256:335661c6e36b5307fbfa9b8b89acf17e0101b130196abeb3b96fd16f717cc0fe