Pith. sign in

Paper Citation Record · LEDGER

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

As of 20 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 42 inbound Pith citation observations for arXiv:2501.13926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13926 v2

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:32:46.670055Z

measured 133 of 133 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:16:31.747844Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.175257Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c96e86b-f010-49e0-be11-62ba37de9d69 · outbound

This paper cites Let's Verify Step by Step.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Let's Verify Step by Step

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.099916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.099916Z digest=sha256:164c3f1671032abacb551ef5f2d2a8f2c5b480aa5972f8131a37718d94e9f016

Observation c2af53d6-7bb8-4a94-adc6-189188921310 · outbound

This paper cites Advances in Neural Information Processing Systems 35, 24824– 24837 (2022).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in Neural Information Processing Systems 35, 24824– 24837 (2022)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.107637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.107637Z digest=sha256:09b83a344caf68f6f679af959ada4b486b13fec623cd7432a4e54a9fdd591379

Observation 5b1264ac-5d5b-41e1-a0d0-73fdd1b5af73 · outbound

This paper cites an unresolved cited work.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.114479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.114479Z digest=sha256:981ac28fa00fbcea5e20006ca65f9135c05c41f78e8ecef104403f0fdd1c0c25

Observation 58823435-7a69-44d3-bdd5-ca5b88c71603 · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.119385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.119385Z digest=sha256:cc90f83e25bff4dc8bb9533658e9a3fa76e077b0dcf94a7981c6c4eb1c8fc54b

Observation d715edc7-9054-449b-9b39-ba828e275aa0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step LLaMA: Open and Efficient Foundation Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.125647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.125647Z digest=sha256:34a972fd1c0438364f0fc7fbfafcea858ff16772ac406fff80f9b82a552935f9

Observation fa23aeed-03d3-4601-8372-55be656dba4d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.131039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.131039Z digest=sha256:25bead91614679e5070c210bd4053cbc4d210bf53a0428a0e12818a71fb9d89e

Observation 10ba5a3e-7c1b-4b64-b42c-67e984a7e63e · outbound

This paper cites : Language models are few-shot learners.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step : Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.136418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.136418Z digest=sha256:053276af2a31dcdf8b0523765bb926a3e6398b66623a0ebf459e9a2becc6cfed

Observation 4425fdaf-1360-4787-87e7-9bfec824f986 · outbound

This paper cites https://openai.com/research/ gpt-4v-system-card.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://openai.com/research/ gpt-4v-system-card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.143412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.143412Z digest=sha256:f73a797fba94a94c4745b5b15e73c2e942a8e6e7336ded288fbfcae8ed62eb36

Observation 326caaab-53ed-4ada-8b17-7716fc141ba0 · outbound

This paper cites In: The Twelfth International Con- ference on Learning Representations (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: The Twelfth International Con- ference on Learning Representations (2024)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.148880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.148880Z digest=sha256:5522d194c639f1e7e8f8201e7316876be88487d3179ea5eaaa9db686ae20f9af

Observation e4fa8aaa-6ff3-48a3-8e67-86dcbf5eb2e7 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.154470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.154470Z digest=sha256:39dc24b6ac6fddcd32d5d1439b2dbc23155de20d9811fc444b568b2d2b443b96

Observation 21664e14-e84f-41bc-9656-ddaaebd19ea2 · outbound

This paper cites https://github.com/open-compass/ opencompass (2023).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://github.com/open-compass/ opencompass (2023)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.159611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.159611Z digest=sha256:6524d8f5aaf1e545c7370880f0a59489fd756372028275ac7aec066a7ac022c3

Observation fe984375-daa3-48bd-8d2e-fb1f890cecf3 · outbound

This paper cites https://chat.openai.com (2023).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://chat.openai.com (2023)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.164810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.164810Z digest=sha256:fb59393d1d8e6acaad91495e1bdec051ce949dc9f68a11bcfc69d4c1254c9119

Observation c6ae4786-356b-4f89-90e3-2a3cf7bc8fbe · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.170631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.170631Z digest=sha256:85bbd88b49eda59c8c1d8112c60dbb179057c5e811c798b14a5288e09eaa5218

Observation dc0e5935-bd31-446b-9a11-3ed715c46f6f · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step MMBench: Is Your Multi-modal Model an All-around Player?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.176609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.176609Z digest=sha256:27953cde355b76d1c3173eb6d335e015e2e3aed32aa8bebe6f9b9da5ae65b1ad

Observation 47fdc94e-a1af-4865-b1cb-74b6ae89b56d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.182807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.182807Z digest=sha256:7a52dcae85d9f98e3595825e1bdbcc0a686959e1bd966d1ecf5f099c80210914

Observation 0041cbb0-3cf1-4b1f-9456-8f02d07b37e6 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step ImageBind-LLM: Multi-modality Instruction Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.188877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.188877Z digest=sha256:8c62030d547e255966550d31bf3f26fb8c9e01e421b9caae2f94b18e18ccc340

Observation 6d5d7674-a9f3-4190-b90c-9bcdb95d3c78 · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.195022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.195022Z digest=sha256:f861eaf8fb5dc7222eaa4bc789c9e527ab5fad5404663790087426e1dbdc608a

Observation d0d93621-ff99-4e46-ac45-d9913ce53f8c · outbound

This paper cites SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.200811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.200811Z digest=sha256:f67e78cf565265cfa6d23da579c2b957a19706e771298d1dea664dd2bf738514

Observation 6a7eecf8-a246-4704-8e63-be92f9a99abf · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.207154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.207154Z digest=sha256:f0f77980e7c473592e342bc7cbbe203ea863f0fac30f4f0cba0a36a593b31d8e

Observation d5ad7b38-04d4-4705-98d0-c893915a6955 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Multimodal Chain-of-Thought Reasoning in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.212962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.212962Z digest=sha256:a55b6f1770df2ce254e81fc024dd815a72e15e40df114964747a272199974905

Observation 12f079a4-6739-44bd-878d-a4f3354393f9 · outbound

This paper cites MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.219119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.219119Z digest=sha256:6c1684a4e4d3941cbbc843e05033a2780be4ae412fea65e5c836f2387df3bbce

Observation fec5fb1f-0d80-48d3-9803-c6b332e0cb0c · outbound

This paper cites [Online].

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step [Online]

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.226490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.226490Z digest=sha256:f07cab4efccc860a398bade606c38c3926ae7c2ca79b805ca5c199e7237da3e4

Observation 989c7960-31cd-46c0-9392-da436cf3c7b3 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.231628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.231628Z digest=sha256:bce7e139e82f7d984beb9ef27832aed34bd40d7ae04d9aefee8773f5a9ea67da

Observation cbc0938e-0cfc-4806-bd53-3cbe8e4af1cf · outbound

This paper cites https://sciverse-cuhk.github.io (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://sciverse-cuhk.github.io (2024)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.236682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.236682Z digest=sha256:fc468cafc5372b83ce6b279e2409173ce55200a0651ae1b975396bba9a228a33

Observation 770f26cd-4e7b-420a-9302-2d62e8913be0 · outbound

This paper cites International Journal on Digital Libraries 23(3), 289–301 (2022).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step International Journal on Digital Libraries 23(3), 289–301 (2022)

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.249799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.249799Z digest=sha256:86ccd8fbaad5fcd6bd87cf093e42da89895335075b5ae9f7908c8cf4487b19f6

Observation 3c790baa-3cfd-493c-894b-011cc57e521c · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.255826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.255826Z digest=sha256:029483198911676a87ed429a1aa21f3be4eebe8759a969e7ad33cb2be4ba4a2b

Observation 38e3d0af-e775-40eb-85b1-2b3015565f22 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.261672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.261672Z digest=sha256:59b1dd566f62acf149acfdbfd84db2fafee492054db67a7b073910f1b5044d1e

Observation 83ab593a-2b36-4a53-99d5-cbc51f927e03 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.266755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.266755Z digest=sha256:1aed632405d264814f68d88410ee35b3cffc0a766089a50262214626502e8e14

Observation 1a2e9baf-e9e4-497e-83af-dcf948d9ed67 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.272390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.272390Z digest=sha256:234c7d63de5d61aedf22ee44ad5640061534455698bd6d7be8491190fb6888f4

Observation 84649c06-02eb-42c6-be3a-b353774fbe7c · outbound

This paper cites BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.277619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.277619Z digest=sha256:e35ce94e6bdffd411f80b8e2e64580b4088a96740b1109cc998edc1e41170190

Observation f4aab891-4530-402d-8c78-58c3679b272b · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Emu3: Next-Token Prediction is All You Need

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.283363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.283363Z digest=sha256:6396e30dabd1a7bbd8504152d1ae1d854f1b35b3ff0d039ce2dcf1740707e356

Observation fba80be6-eedd-461f-93ee-9997a6459062 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.289517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.289517Z digest=sha256:4aed7a81a921deacfc50a1880223b808d43b58c37598f42f9c5927920d0b3758

Observation 107a212a-197f-4078-8c6e-a791668d6558 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.295785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.295785Z digest=sha256:9c498cb20411b04866e04fb0f477daa800b4f9f25c5afa83752bc17f04df6798

Observation 3d4b80b3-9e4a-4396-89de-ef5e58d877a8 · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.169018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.300218Z digest=sha256:1fc4dd6b3d628350010edcfc69c1a547215bf52f044951c740ad7ec8d8824298

Observation 5e8bef2b-c452-401f-8e51-814fde43fc4e · outbound

This paper cites Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.304329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.304329Z digest=sha256:eb64bd305d571552d9939631f7c70f2b0c16607c7928a71f82e447395db251f3

Observation b3f98590-752b-423a-adb8-a0d733587b04 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Teaching Large Language Models to Reason with Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.309279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.309279Z digest=sha256:7aa8c71fe85e135b3ff0255aa60dc661be03bd9e26d9e88e9a97478df32ac6ff

Observation be8edc1b-2d84-466a-aeca-f54a9f5890cb · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.315309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.315309Z digest=sha256:4d906d2021e73864f6b3504aa4ab6fd996a99d94225d976216e46742e227b630

Observation 4f1427ae-6884-4a09-bc76-0d79c12e2649 · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in Neural Information Processing Systems 36 (2024)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.149499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.319898Z digest=sha256:f984970f8b07bc8c96c1a1df5d51dd47278c09a11b3b58bc0a5f46332d27fd3a

Observation 18e7c0ae-9925-4d1b-a088-169b569ebbf9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.324481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.324481Z digest=sha256:408f61dbc32702d5e8ba8926473b0f9c5dd2b37ed8693a3b64f5a7651f5830fd

Observation 09ee2ccc-5b99-4313-9c42-3350cd6d931f · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.337555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.337555Z digest=sha256:c2d9960e61b80e97afcf768e85708acf1b3281a566f8ff14b1f6cf96f84b5aa1

Observation 73b2a888-cbde-4897-8090-6905b21c5e3a · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.345905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.345905Z digest=sha256:4eef7d5a808862a1d17c272738725254c2e265bb40bb8412ed0908fd3184260a

Observation bed49623-bcc8-4193-8992-de02fa411719 · outbound

This paper cites DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.351972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.351972Z digest=sha256:915becf203f7692ea097f965e4eba8e7072c11b3ec8f949dec0deb8f63d1427d

Observation 2ad8f5d2-9f3d-4e39-bd3b-bef64c88cc7e · outbound

This paper cites Make Every Move Count: LLM-based High-Quality RTL Code Generation Using MCTS.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Make Every Move Count: LLM-based High-Quality RTL Code Generation Using MCTS

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.358123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.358123Z digest=sha256:1e1fef1816bb10265c2df1d3fdc7f45ec65057fa53d933db98bf870beaa165b6

Observation bc5eaa5d-9552-46bc-b1f1-ad47bd0a39a4 · outbound

This paper cites In: International Conference on Machine Learning, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: International Conference on Machine Learning, pp

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.134016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.363676Z digest=sha256:e868490f593a0abcefc813329216679a4389fc125e3360776fa0e12afddc8d8c

Observation 64964ef5-ccb0-4393-9ac8-e7a758145b0f · outbound

This paper cites AFlow: Automating Agentic Workflow Generation.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step AFlow: Automating Agentic Workflow Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.368693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.368693Z digest=sha256:12743028b9745b360616e0f7f48b8aba4eade0a3ed65397f393677be41595c26

Observation 7ce4e38b-3bb0-459d-baa7-0fb85eb0556c · outbound

This paper cites ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.377436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.377436Z digest=sha256:d6de772957b467b939c1d87430cc6e6bb90b16bd46fd5f11c495c80af6b9c522

Observation 6f0fe4ad-7de2-45a4-bc9b-7cb90ad313d2 · outbound

This paper cites arXiv preprint arXiv:2409.01392 (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step arXiv preprint arXiv:2409.01392 (2024)

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.383805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.383805Z digest=sha256:eebe9b8664ddbefcfc9f18e5fcf3793c09bd6ba591f3203ee697670f512f4c6d

Observation 29c6c530-c439-44ee-8e90-2d1218ebe5a9 · outbound

This paper cites Advances in neural infor- mation processing systems 35, 22199–22213 (2022).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in neural infor- mation processing systems 35, 22199–22213 (2022)

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.119198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.389252Z digest=sha256:2e067c7f3961b72f59b080978499dff5b8f47d3f51f859d40809059e49b9d172

Observation 245fd680-d181-4e80-8d3c-104f62301a70 · outbound

This paper cites Chain-of-Thought Reasoning Without Prompting.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Chain-of-Thought Reasoning Without Prompting

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.395722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.395722Z digest=sha256:2a87e6ac8af23ee92643ac7991e082b60ca3eb1077d673e4179325f71d1d69bb

Observation 9550f5ca-5715-461a-b489-d4be99ff68a2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Training Verifiers to Solve Math Word Problems

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.401062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.401062Z digest=sha256:4845e1be3121221ff0239b7ab023de265267e895cd98b945ecd7ab676fcb9c87

Observation abe25246-05e8-49ed-8ab2-8a362b467a93 · outbound

This paper cites https://openai.com/index/ learning-to-reason-with-llms/.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://openai.com/index/ learning-to-reason-with-llms/

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.103117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.406595Z digest=sha256:aedfa090ef952b79c9c979056b1691f95124c94b6cde2c73934ce2c7e8b46366

Observation 70124feb-1e7e-4d0b-b5b2-1f4228a692de · outbound

This paper cites Advances in neural information processing systems 30 (2017).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in neural information processing systems 30 (2017)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.086408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.413534Z digest=sha256:d356ca050c354ac1e456e5a89d2009b20dec6331fdcb3f4a3ceb3fadc977ecc7

Observation 141cfdcd-5520-4890-8ba0-50dcfe5e650e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Constitutional AI: Harmlessness from AI Feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.419501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.419501Z digest=sha256:1ae1bbb644e0fbbf422a90126fdfbe0b0addf98ca0e3b62ccff3bb55bae0e5c4

Observation aa0251be-08c1-431c-9df8-e8cee779af96 · outbound

This paper cites Robotics Research: Volume 1, 161–176 (2018).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Robotics Research: Volume 1, 161–176 (2018)

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.068969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.431765Z digest=sha256:c98e16b41321837b8254ba4643cc63e89e651ad447a853d2a4cb12afc56fc2fa

Observation dd549d22-6eaf-4d51-b6c9-afd1c0b4758b · outbound

This paper cites Dueling RL: Reinforcement Learning with Trajectory Preferences.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.438395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.438395Z digest=sha256:136626c910ab640ef5fbcfe9fe724d2088114590205075938e68d3133bce62ae

Observation 0bc97bd7-3bd9-4cc1-a4cc-ce7d25c1542d · outbound

This paper cites Advances in neural information processing systems 26 (2013).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in neural information processing systems 26 (2013)

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.052863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.444302Z digest=sha256:c52c9be1ce687bdbbd283637c6c41d0e590a02da35cdc17148b59d64088a948f

Observation fe872f6b-218b-4fbf-b8bb-544150c59223 · outbound

This paper cites Machine learning 97, 327–351 (2014).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Machine learning 97, 327–351 (2014)

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.037129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.451017Z digest=sha256:de3259487981a21b30efa37c5f276873bfc2c3bc56a16f0149af4b663c0b9c04

Observation 487da223-b5cd-4c8b-b54b-6a13d2a10d39 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.457214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.457214Z digest=sha256:f7de6f3303ab9611fe7a3f147ed3f9795161052691733d7d634a6acb53c1e514

Observation a030b60d-d2cf-4e86-8157-8129fb437613 · outbound

This paper cites the method of paired comparisons.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step the method of paired comparisons

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.012302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.462981Z digest=sha256:f711cae2a60e6d59f8f719c0f458ff308900dd9f721c639febc18c13c3bae7f3

Observation f27b1868-52a3-4671-8ce8-c4b67b1c9b62 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Proximal Policy Optimization Algorithms

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.468470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.468470Z digest=sha256:20e628509944aaebd4001c84ecb7a3eea021cd3216ddc9a1e3a03a3d54f12a20

Observation 06a6ae0c-116b-420e-b4ca-318d7bb2800a · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in Neural Information Processing Systems 36 (2024)

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.993699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.474113Z digest=sha256:bc34a5892d3c06334dac2b991e903b0f0aecfd118f6ffac7f70a9bc184aa933b

Observation d2d891ec-a4d2-468d-b4d0-980f3d974430 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.479044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.479044Z digest=sha256:9eb03306ac710733a78fd2ad42552600cba36baac17a302e5a4a8a4fcc84e1f7

Observation 1e07bc7b-13be-42f2-a374-0c0c02c58443 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.484359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.484359Z digest=sha256:46a2ac85fe11973570c19877f1e394c4a38e366208571842a495d2df5682f0b4

Observation ee2a90fd-5316-4a53-beeb-f059953d7a95 · outbound

This paper cites Aligning CodeLLMs with Direct Preference Optimization.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Aligning CodeLLMs with Direct Preference Optimization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.490134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.490134Z digest=sha256:aa4ab50f03ec536006e5cf510c59a2aff742907e7ddc51a080c920ae0639d348

Observation 4bf62b0e-ded0-4ade-8c5b-5219cb248b25 · outbound

This paper cites Code-Optimise: Self-Generated Preference Data for Correctness and Efficiency.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Code-Optimise: Self-Generated Preference Data for Correctness and Efficiency

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.495881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.495881Z digest=sha256:0b38d8f2184f17ce8343ed9b01c09f87a1b8473adc3cc55def2f07d3e6fb2062

Observation 6d6e59f7-a3ac-4e3e-8d71-4650e0e8be58 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Improve Vision Language Model Chain-of-thought Reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.501314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.501314Z digest=sha256:a5cd8ffc1e50859f6f502b2cfb7671da08f4da880558e55e32e9d5aea26e06dc

Observation 00e4ef6e-9293-4a0b-aa0c-5a4531ba478a · outbound

This paper cites https://github.com/meta-llama/llama3/ blob/main/MODEL CARD.md.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://github.com/meta-llama/llama3/ blob/main/MODEL CARD.md

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.975412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.506642Z digest=sha256:f437561da43f02516d771c82015dd0f43e22b68bb8f274e3b1955460573a5e9c

Observation c85d5cab-dafa-44af-8aee-e51bfa7aad92 · outbound

This paper cites https://openai.com/ index/hello-gpt-4o/ (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://openai.com/ index/hello-gpt-4o/ (2024)

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.958500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.511772Z digest=sha256:c3b58b4264631c8976b62e7a6e64ac9cd9c629284d456d1d65f10b89e82daa68

Observation 9a87b054-f017-4058-a30f-7e36168e808c · outbound

This paper cites GPT-4 Technical Report.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step GPT-4 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.517639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.517639Z digest=sha256:5dd49de9a7f351f6c9570840168faa107e54da4e61d2c21894105f737db9c95c

Observation be0612c3-4cfe-41d9-bd7f-92ff9e86d395 · outbound

This paper cites In: International Conference on Machine Learn- ing, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: International Conference on Machine Learn- ing, pp

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.942121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.522937Z digest=sha256:5f731695dec375d51570371609418e1c5130139ff03a662f1617ce342b176fbe

Observation 072dc709-42dc-48c9-9b4b-eafc430dfd08 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.527690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.527690Z digest=sha256:a533095bc3f95eb0097365f2505fad1499a41958029ddb052a02f890820bbf52

Observation 50d7ee21-67eb-47af-a8c7-c6240e8089d9 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.532426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.532426Z digest=sha256:7c52610624e6d6dd5fe3de4e9f1d9307c093c748df8a75483da5953ddbebd720

Observation dfea0120-e7eb-4e71-bf51-1c91556f40bd · outbound

This paper cites In: The Twelfth International Con- ference on Learning Representations (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: The Twelfth International Con- ference on Learning Representations (2024)

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.925393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.538711Z digest=sha256:96b4914fd66e8ac2da3a6301313a6014c5eb5f22437d2d897e40c3f155ebf22c

Observation 8a630b36-2620-440c-97a1-2abbec01dcf2 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.544049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.544049Z digest=sha256:1deea37d508b6158cb38cbe93bed88bbacf0a66bc144585465ed40df1cf91a4b

Observation ce03faab-d2cf-4f22-9a0b-445db9186bb3 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.550031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.550031Z digest=sha256:7cd2427fa16ad9116c577244133d8b099be06d5b7a7cb19863c041640cf3fe18

Observation e3af367b-9dda-49f8-88e6-d3f8d126f804 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.554937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.554937Z digest=sha256:5d466820abf8ab34b3f580d2537e4725bc1dc7c103bdf7dfef6afca8a2f7b479

Observation 1f6435dc-bb0c-4923-9fe9-450e76224116 · outbound

This paper cites CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.561266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.561266Z digest=sha256:3f986d07fa36868386ea0231fdf10341be966d21feb2852a989e0f00733e6007

Observation 8ca7e623-3402-4938-8e5d-de2252e40115 · outbound

This paper cites Advances in neural information processing systems 30 (2017).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in neural information processing systems 30 (2017)

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.906998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.567076Z digest=sha256:a44188cdf9279f8c11423b69ded4b3c84996b97f26a58372986c07af16d8916e

Observation cd032d3d-c657-49ad-b66f-707e0762545e · outbound

This paper cites an unresolved cited work.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:32:47.890749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.572873Z digest=sha256:e668513c01426e857cfb447c1549aceb23fc007db445b78b4ed4c96a8e8ca9f0

Observation f935dd1b-0b4e-436a-9b2b-46624a795abb · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.580857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.580857Z digest=sha256:5ead8b300f314035538b8bd33a393d1fd53885be505bbee0e449b814873974f6

Observation affa5a69-a4cb-415c-89b4-cca260047913 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.863859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T15:32:46.589733Z digest=sha256:186c5fca13ed7245ea7149b9c42538ce83a6a473d6d2b7ba7d2d86e339caae6f

Observation 4668aa73-4b97-4cd5-9c65-150ab26f9adb · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Autoregressive Image Generation without Vector Quantization

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.600985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.600985Z digest=sha256:e5c9c68e3326ca33e7e351b86428b5debe85e595107bfa38e5921532b10a5224

Observation c164f945-7a93-4d79-9fd0-a2960160d058 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.617205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.617205Z digest=sha256:96a7995a10722fa52b292c7d4ebdb820580e1d11c00fef7e2b123f70cae8fa5f

Observation 23f23d38-ebe5-4dc7-b7b4-7dd8d7d9b13e · outbound

This paper cites Iterative Reasoning Preference Optimization.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Iterative Reasoning Preference Optimization

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.622548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.622548Z digest=sha256:71e2460d8c679c3e882e4d3d02970be4dc271fac3f225354a1c0fd7dbfaa1fb7

Observation ff62fa16-7ee7-4a7f-95ab-c2160dd40716 · outbound

This paper cites Qwen2 Technical Report.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Qwen2 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.629650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.629650Z digest=sha256:bb79cc77325e3a847f4df7d5e1f7d8e67f634db6f07be6e4455c6ec2452a7ad8

Observation 4e40192f-9fe9-4985-93f9-d080ab23c06e · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step High-Resolution Image Synthesis with Latent Diffusion Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.637932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.637932Z digest=sha256:9bfbc5f8482a639491f04ca7996ebb1527dfef39eb7e1ff7d959059c4de7eb5a

Observation 5a690ba6-c7be-4081-b376-ac467472254e · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.646084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.646084Z digest=sha256:1f1d202b50380a98040648003975935be7aa9c444d2a1b1909973b3936770445

Observation 5c44f91b-121b-4bc5-a07f-7a3947dbb51d · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.654313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.654313Z digest=sha256:c8bcb72084d23a4ad80e81c67c35ef751d21bbd0837b93974f9e592a0212831b

Observation f934ddae-9fec-4718-9936-829a27118b9a · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.660605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.660605Z digest=sha256:1eaf274b7ee61b8cc53a777d5d6bc77a607da1c8b3bf088ac0483935666bb615

Observation 1356bd90-dd7b-4f44-ac42-bb7ae7c803a0 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Training Language Models to Self-Correct via Reinforcement Learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.665734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.665734Z digest=sha256:62510b02aa35ee98997e76cb79c3f725132c9c897ce693ff3d357c9bb23323c9

Observation 442640bb-bc79-4d78-8cda-1086fa7fb338 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Large Language Models Cannot Self-Correct Reasoning Yet

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.670055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.670055Z digest=sha256:d4be638b0e3027856bbe9307494f9fe6dd50a5c3221c3c152b4d9636c1a3d091

Pith citing papers

Observation e42c13fd-978a-4186-a935-9deb61109a38 · inbound

Multimodal LLMs Can Reason about Aesthetics in Zero-Shot cites this paper.

Multimodal LLMs Can Reason about Aesthetics in Zero-Shot Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:16:09.308555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:16:09.308555Z digest=sha256:9621d1323efaab1da4c68f60f2de9034fa262d7b1ff1f40087e1106e1a009adc

Observation 48f17002-9e99-4caf-acf4-ae4711e6c14c · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 234

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.448367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:4a48ac3607ba2eab56ca4eae248e6234e4d54805963b95b60d5b903ec344080d

Observation fba5499c-c290-4ca8-a6e6-37db20ab6514 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.651274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:03f0c1b7b1d887cc2129114e9fb165173fae33f1a6ae8f627a8f4a87aaa8e39f

Observation 9484a144-9274-463c-82ce-9991b15d12af · inbound

$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark cites this paper.

$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:16:31.747844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:16:31.747844Z digest=sha256:37238106738d397b17df3065d340f2219369c8cb903941ecae250d545409ef62

Observation 9ac288c0-7b3c-4f6e-bd9f-4259ee2f26e4 · inbound

Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing cites this paper.

Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:11:32.527138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:11:32.527138Z digest=sha256:96d9131c52b3c50d09cd1d416b4481337c572a7af60f5f695309e18af8df65c0

Observation ecc3ea90-68f6-47bd-940c-49c6ef7f5b02 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.021279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.021279Z digest=sha256:123597d229f6a1723f7a683afd5d88cd6c386d62cc0084e851d9f383cb8591e5

Observation 73905dbd-0e28-4e4a-8db7-a7e752140022 · inbound

From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning cites this paper.

From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:34.890607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:34.890607Z digest=sha256:542fb1650ac93fce5f81d5b25094e2ec8fdd4eecae2de68460248c421d8057d0

Observation 38867dd6-2c10-4a26-9831-075800eaffb6 · inbound

T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT cites this paper.

T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:41:47.575037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:41:47.575037Z digest=sha256:87bdc7f4e6752c7dfc45c4b03fadfad4b16d2b33f6ba9789ba7e1b8c801cd449

Observation c6f9f267-9f74-4ea7-9ee7-83052f813b8b · inbound

DanceGRPO: Unleashing GRPO on Visual Generation cites this paper.

DanceGRPO: Unleashing GRPO on Visual Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:28:26.031397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T22:28:24.929046Z digest=sha256:d3f8d56a0ec1d5af3623f802f389bb0e438255a8f41858e26cb1d19249018d92

Observation 7528f7ae-75b6-40a4-bb11-02e6e231d17d · inbound

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO cites this paper.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.175229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.175229Z digest=sha256:95b55ed18d33c3b488b3aab558986fa84a627312263f9673cbbf70576eba5598

Observation 1fc057b5-f3db-4e60-b51b-77bf28e69907 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.137002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.137002Z digest=sha256:134bbfdbfd8a22cc1a12fc560747e72c6c8e86b8ccd3889e8f5acad6898366a8

Observation 052310a6-09eb-4896-a7dd-72062cfa63af · inbound

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO cites this paper.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.520817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.520817Z digest=sha256:db828917df61c1334c84f17a524faa28cfab8c4b59965486679f02212f399053

Observation d2ea6fd7-5eff-4b53-b889-422edfe5faf5 · inbound

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning cites this paper.

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:03.001071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:03.001071Z digest=sha256:5d152fe4f568060c4b6f8247df2211c49b755109e85cb37fc1112754509cd043

Observation 051c8f17-2752-4dd2-86b8-fe78be8377d2 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.310672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.310672Z digest=sha256:6426449bf92898b13abba2b71f352df241b75d50bdbfe40511d4246498e4e174

Observation 5c976fdc-86e1-4356-bcef-5e81b52d8b8e · inbound

Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation cites this paper.

Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:36.335219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:36.335219Z digest=sha256:fd3e90a638d21af49fa4716c2fb9d36a2df13d9e015b03f996245014a8bfb544

Observation 3e8053ea-4c2d-445e-aebf-e7b2b1ee8a09 · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.895103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.895103Z digest=sha256:a145021e6e91dc7d0f1845a4c72b239500085489a473477f6e8e44db53c2a4d5

Observation c30ed327-5349-46e7-b27a-bb7d9de92dad · inbound

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation cites this paper.

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.044539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.044539Z digest=sha256:d7bbfb0b059bc2e3a60416ffe2cb64b4234804f1677ce5906a974af4138d6d1f

Observation 8a0adae8-f90e-48c0-a83b-bca2917ef604 · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.740034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.740034Z digest=sha256:5205f458b37154c35858ff0830cb56ac9944550ed6a0ea58b0bdaf86345f87a3

Observation b3f0c90b-2bef-45b1-972f-8a68e39d3793 · inbound

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation cites this paper.

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:32.622198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:32.622198Z digest=sha256:623d89b3d8e56856e9089969787eb5557450e53b51cfbcd6e052e49516329152

Observation 081cfe8c-5871-4cda-a8be-8a5bfe6693bf · inbound

ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL cites this paper.

ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:21.683953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:21.683953Z digest=sha256:98a7bd5ea8b67d8dae58f08a7f59ee30b42b987ff139878132db0d55def27930

Observation 553d2b69-b567-4d94-9e80-6517d6fa6cf0 · inbound

TIIF-Bench: How Does Your T2I Model Follow Your Instructions? cites this paper.

TIIF-Bench: How Does Your T2I Model Follow Your Instructions? Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:20.782033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:20.782033Z digest=sha256:a5773f866c7fa0d759da5a6a892cdba69475c19f0c4d86b6db8aac2aa1349082

Observation f932e840-318a-45e1-9f6f-875b17f5d0a0 · inbound

How Far Are We from Generating Missing Modalities with Foundation Models? cites this paper.

How Far Are We from Generating Missing Modalities with Foundation Models? Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:15:33.584502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T08:15:12.947854Z digest=sha256:8035d2a1b253a44efb6a6f5cc58c6b57dcc8c2293a2a74f136d0f210c515a2fe

Observation d4c71618-6fcb-43ad-bb39-3eb3b7ef5313 · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:48.738793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:48.738793Z digest=sha256:ad2f6eaf60a08302f4e6c69cb8970068e1713be5e6a38c4f7fd5a944f3d4ff1b

Observation 23dca6e3-ebf9-46de-8494-cb092f44b9b5 · inbound

Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment cites this paper.

Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:22:44.051500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:22:44.051500Z digest=sha256:264cc14ae4219f81b47658390cfb31ade7242939c36711170bb547f748b11412

Observation 96b75dea-f716-4ad6-8b6c-2b525eb43572 · inbound

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL cites this paper.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.511095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.511095Z digest=sha256:5bc164961fa2b28f5b02782857754434fe131adbcf6b09fff4277c3c3556c9be

Observation 294919e4-d9e3-4c01-9070-ad012af279f8 · inbound

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies cites this paper.

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:20.308003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:20.308003Z digest=sha256:71a11c211c29eccde05348289b6ef7daedb513988cd68fc8261e3397cc6fc58e

Observation 26c2409b-046a-4e3e-87f2-a4338ad6850b · inbound

RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought cites this paper.

RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:33:02.188477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T08:32:20.566798Z digest=sha256:6b04488e3daedd67f07021644bc9a94e8c1df103c9bf084addafc3cd4c0d9a4e

Observation b8fadb8b-3881-4257-9666-e71df3074dc1 · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.744905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:b94236872c15787e24cad6e6f6db4df23ac7a0aaf1fbc2f9ee438a08f85ea2b0

Observation 1255e509-c867-452e-a7d0-a6c4e527b6b2 · inbound

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation cites this paper.

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:42.906185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:42.906185Z digest=sha256:26a5094448bd3485e40f8d2eaacd16f24bf3c037e322d3d3d1a189ad053c6bd9

Observation 4aeb23cd-e3eb-4a93-9403-d66de36d0679 · inbound

Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling cites this paper.

Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T18:22:42.160153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:22:42.160153Z digest=sha256:62094905b72a7603f9dc5e70d693704bf73c44a9d11526c8a2b3a5897fd6bcff

Observation e73c98f9-504a-442d-99f6-92ef727bdfd2 · inbound

MultiRef: Controllable Image Generation with Multiple Visual References cites this paper.

MultiRef: Controllable Image Generation with Multiple Visual References Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T22:32:51.804152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:32:51.804152Z digest=sha256:55c72fd8ae6343b72f2bcdce7dd7bc1b9dc28af9922d126c3e4b1e76e48e3307

Observation 1c41b22a-8423-4dd8-9ff9-b14b3a0b11d8 · inbound

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation cites this paper.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.432203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.432203Z digest=sha256:4f4debb012e3770a2a2522eb074f8fbb6dcdf870f2704215cccd6eff1c372cd9

Observation c7d72a8c-5a3b-47d9-8018-1bc7d72d3564 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:07.963130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:07.963130Z digest=sha256:f45bfcb733f2893f1aabf01107e599b62e9642cf31587314e1a05be630d607ed

Observation 23d37186-de7a-4292-9018-bf899153e531 · inbound

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition cites this paper.

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.343091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T00:20:58.483350Z digest=sha256:bba8f21414ed076a6eae4a1d48e6eb7e079bf90c630fcea350df87ca51f2cbbc

Observation d3c4c922-b22d-4177-b0c6-590127eb25af · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:847f94c0e98c6db23df6180b162a368b1c59eeaf1272103277fed7ee62f80746

Observation 17c05c7b-cb86-4bb6-a657-19a2bcbeb692 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.197765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.197765Z digest=sha256:7536a4ab9dc6a89998ef3c9be4facf7c0148ff5d2f0afcf6080df4dafbec6c1f

Observation 53b29446-6ab5-4120-bea8-1451dcd9a279 · inbound

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation cites this paper.

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:50:37.069714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T12:50:13.764159Z digest=sha256:746eb807b6fd6ef111f66dbd60d4aad6a9a2b256b96cfcc6ac24b4b2b87b4f0d

Observation aafb7341-2c9b-4aa3-a2c7-b6e392123764 · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.520773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:71c848b29237c0cac6d772f61c434df9b669adeaf9a9f0a79e5faf131b913e37

Observation 1ea53590-1d6f-4963-b40a-5219ead2ee88 · inbound

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning cites this paper.

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:25.295363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T19:36:57.231932Z digest=sha256:eae1952abc41e3299269abbaca715e6e70fd34b041fa5081f567a7e30b440338

Observation e823c830-b45d-49d7-92a8-b36f8676a353 · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 101

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.837081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:63ade2a68caae334051191bb0291d45a0db8e54082beb483797b77ad5e0b7b1c

Observation fdf1b1a6-3013-4842-8d7f-e210294a46c2 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:58.177090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:ef4cef406e51ec1ed75f45e4c5fb62e3d9f9a2e9c0246cef13c0fadec65c41e4

Observation 1730a604-eb8c-49ea-9e3f-50043610d767 · inbound

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning cites this paper.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:78ff6dd8ba3fb0fb630144015e6375d45beaf4d4e3609945e3881d19a99ab8b1