Pith. sign in

Paper Citation Record · LEDGER

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

As of 20 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 10 inbound Pith citation observations for arXiv:2505.13031.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13031 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:26:58.409542Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:10.990159Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:06:29.440180Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7431ca4d-00fc-4ac9-ba5b-c7a196d5b8f2 · outbound

This paper cites Qwen2.5-VL Technical Report.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.092350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.092350Z digest=sha256:83af90e4d18a0dbf44ea1c84f97731a1cad9c280048ee171e91682f3e82d8f3a

Observation a51e9c89-782f-4e8b-966d-59c6fb7bd70c · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.099262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.099262Z digest=sha256:8acb311d4f4bb4eb035d456ac157607a5bb65c6987362865be9d0822a5249cbb

Observation 389c266b-d5d5-4b08-8564-f988b4986220 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.105088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.105088Z digest=sha256:56806cf29d853a40ddfffdf2e659a6d49fc22c5e28d450d90c1a48c2dc7cead2

Observation 1b0cfad7-a596-43a5-9889-8a122e251c33 · outbound

This paper cites r1-v: Reinforcing super generalization ability in vision-language models with less than 3.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO r1-v: Reinforcing super generalization ability in vision-language models with less than 3

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.518115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:26:58.110901Z digest=sha256:b36b921b03ca766f5bf4a0b134bf55aa277b3e496c8ce08d69b133681390c8a1

Observation 8d82be19-ace8-4e34-a7fa-675d7a3ffa2d · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.116559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.116559Z digest=sha256:e6c03452517a506ba3bd2fa3f2780292f56531d5212a72dc57dde7dbff6ddf6a

Observation 766af505-288d-4f51-95e5-2adbc908af8d · outbound

This paper cites Science China Information Sciences67(12), 220101 (2024) 10.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Science China Information Sciences67(12), 220101 (2024) 10

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.488898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:26:58.121570Z digest=sha256:2b05100a184acc04f1c1c75a0acddf9803a66c94985e0259f741032ebb43cab0

Observation d9523808-2f9f-45e0-8845-8718a37b2ac6 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.126897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.126897Z digest=sha256:17fcdcb07139a51cc6d9221dadb67401e2cea2e266fab7510f7be4801320e0e5

Observation 83349dd2-4b7d-4060-8d87-1c24a8796109 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.133735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.133735Z digest=sha256:322ff4e6b9f5998fae5535b3abbbbf2e078397c3b0170cb8ae30f5bc1b14affb

Observation 71cab7d0-ea13-4b25-88de-9852fedef3cd · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.139315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.139315Z digest=sha256:75c8f5e5a2861de4a535f2ef6e311a16df063ffbcfc9fdd1aa4be3c634d72a94

Observation eee365db-c203-4bf0-a66f-19a157c902cb · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.146231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.146231Z digest=sha256:e4d5cf2591d7dfef3c18f332b1834288fcc11d0240600a9b934323be2f84a987

Observation 257300d6-e990-4252-a71f-e6d7c8cfda74 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.153083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.153083Z digest=sha256:c1886f140c47630cdaf1e421d62783c78a85e18c45d8aa034cd20736a24697cc

Observation e2edbfb5-4dce-4b66-9659-d94dd2214f9d · outbound

This paper cites Advances in Neural Information Processing Systems pp.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems pp

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.470305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:26:58.158664Z digest=sha256:5da0b139c2125d75b0bb59528b15a675aba284ea8ee9db3ec1ba19133cd58a9f

Observation 25e2075e-38a0-460a-90fe-77141058ab01 · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:26:59.450727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:26:58.163671Z digest=sha256:28f978f303e4f6f4b9bb16379584a0f64945d34cf778df1c5791b1f8bc4156dc

Observation 3fcd54f2-d329-46df-825b-d4fff51a8731 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.170393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.170393Z digest=sha256:6d7bd0087ca4bb94a450ab56d335a71d6c69191c7440ab2ec7beb9123eef9e38

Observation 7528f7ae-75b6-40a4-bb11-02e6e231d17d · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.175229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.175229Z digest=sha256:95b55ed18d33c3b488b3aab558986fa84a627312263f9673cbbf70576eba5598

Observation daf4f96a-731d-42e2-8f39-e4b1da2c1424 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.180483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.180483Z digest=sha256:fc1356a0c8544a2c30954d673d25c777a7d5320e618adca7bda93023b6e35155

Observation a90a0824-2081-47b8-928f-21aa3cabc3d1 · outbound

This paper cites ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.185469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.185469Z digest=sha256:7e7c892263802362fe4cc65f655fe850c00cdac17e5972c037bc0ddca82b4462

Observation 8f7ffe76-90bb-44b7-ad54-10d5e720510a · outbound

This paper cites ARGS: Alignment as Reward-Guided Search.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ARGS: Alignment as Reward-Guided Search

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.192292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.192292Z digest=sha256:1a9f4299a7674112ce7651128174931438054339dffeb034ad1f4a187951fb26

Observation a4882627-c6d9-4a15-b758-8c0764825b81 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Training Language Models to Self-Correct via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.197909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.197909Z digest=sha256:540f7d6d651172cdadb62fced4ad0bfc882eddb4b15ffabf83cbc8ab2ff087f1

Observation 970508f6-7142-4887-a75f-fbc7eeecfd46 · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.203015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.203015Z digest=sha256:95f101a125e1a0c6ca5092aa1f4c6c4f117a34189ad82998f1edbe230e9ff844

Observation 18aca474-96d7-4523-ba18-750605542886 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.208127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.208127Z digest=sha256:6183d3679e0bebee96f0b3cb754f4d01cec6d4c6799ed1b8fab821df70cd27bd

Observation 6a3dfc0b-2b32-446e-a4c0-85130968bbd0 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.214608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.214608Z digest=sha256:f3ebf026a26e91ebafa91f4cc32a62fa170ea7a85d67b5ab41d12a2e37e8416d

Observation 97f83c26-75db-462e-ab39-20c7d35cdd82 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.220542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.220542Z digest=sha256:78b6ed6f9b350906bbafacd0b08c92a4ffac082ae10c850acd5091d80f67fe39

Observation 7598a37a-d303-4a7a-b95d-38b2f95001f9 · outbound

This paper cites https://openai.com/o1 (2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO https://openai.com/o1 (2024)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.420022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:26:58.225798Z digest=sha256:80747493d8886b312bf701c4a62e8bdc3898d5e33d473e7254f62e808ea2ca2e

Observation 1a45b6d6-9c64-4ef2-be54-7642c52af2ef · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:26:59.401571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:26:58.231910Z digest=sha256:d1857a4b7927bb26f05db5afb8796544479e733510d34097ff8215221794dd5d

Observation 291b28c4-8111-430b-9d8c-04cdec535f42 · outbound

This paper cites Transfer between Modalities with MetaQueries.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Transfer between Modalities with MetaQueries

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.237599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.237599Z digest=sha256:2a2817cec99c7de8a0100f0916090315cf1c1a8f80a98787cbf1743183bb1a81

Observation 85580cca-6308-4b8d-8487-64a66c2ff527 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.243052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.243052Z digest=sha256:41fe9838ea595ca34a01fd8fa6389aa839989a5212b559c443ae4f1fd0d3fd47

Observation 919c26c5-2940-41b1-b90f-853ae5be2a05 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.249057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.249057Z digest=sha256:a76444203e5e6edee7140b868f63a5590f7cdf3ef6739800426d3d668bec899a

Observation c2050e8c-61c5-498d-b050-32f0d785b733 · outbound

This paper cites Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.255579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.255579Z digest=sha256:d8d1dc72c50962c46439d93e393ef3e1835684ce709fb030182bdab6dd7cbf00

Observation ec67038c-8d8a-4cdd-b29e-5fc2732a4d94 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.261282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.261282Z digest=sha256:08641f6b2c8e2995bf8608180b0b7eb5691f882602032fbfcbdd4880c5bd2585

Observation 28a7daf4-7488-4264-a553-191935772155 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.267389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.267389Z digest=sha256:4f2532024dc276095b59781fe6001cfcc16d8e36a432627a5ca87e5aa2ea0dca

Observation fd21ddc3-1885-4e10-86b5-078be8500a1c · outbound

This paper cites URL https://laion.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO URL https://laion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.361428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:26:58.272917Z digest=sha256:182ed20f60f91ab2487fee781f3c92370481e77d3d0cb6ec7c873fc669af130b

Observation dc57ee99-1fd9-474c-8474-4f5352a5cc1e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.278116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.278116Z digest=sha256:e9f8689fb8cf12986812abd1ad025ee88bb780155f6c0ebeb328ceb3e5048f3e

Observation 201000fc-ad77-4f4e-bc67-e3e3d979dec3 · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:26:59.340085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:26:58.283274Z digest=sha256:aa4e3531c45ee94061a17ffe778dc66f6db24e13c0a7f1937031e88b893db052

Observation 3c262794-8a07-4f9e-bede-c27c41c66ddf · outbound

This paper cites Advances in Neural Information Processing Systems36(2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems36(2024)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.288057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.288057Z digest=sha256:5b267814a9bc49474282407d761072629caae1f74952c220540784472b3dfee4

Observation 4c803fed-c3c5-4f6d-8871-30da9ed25ac7 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.292902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.292902Z digest=sha256:f190f2a6decd3ebe1fb8a79995bbc270ed4aa993fcf2abf544ed3924ed50bb5c

Observation e4632963-3123-4104-b269-d512bd4d163d · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.298120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.298120Z digest=sha256:cc08357f5b93f526c3ec3997722e992fbc6f48876dbd0e9057ad7c52d4fc5c2c

Observation 98a15ab1-62d7-4d53-9d89-e417288b8476 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emu: Generative Pretraining in Multimodality

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.303151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.303151Z digest=sha256:115dec75040eff817ede31e591efc2303dda35b181534b090ff868bb26783585

Observation c29421bb-db06-4785-8940-54aa1b3a4209 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.308208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.308208Z digest=sha256:dfdd77fe70c556dab8e1305c639ecca2180df360d8d0a61243caff189dd4e678

Observation 66b03bf8-fcd7-4704-bded-6c384a40c4e6 · outbound

This paper cites https://qwenlm.github.io/blog/qwen3/ (2025).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO https://qwenlm.github.io/blog/qwen3/ (2025)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.291607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:26:58.313105Z digest=sha256:81aa1641cb289aa2ae731d0b1ec9a1d5f382a4f648f40232e2d38572a4fbf829

Observation 24f5c3f4-324c-4a48-9a70-8ebbefc79636 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.318304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.318304Z digest=sha256:ac7b4129c9e235cb8101b3a15bf1abf5775abb236de9e5e332fdf6eff2c6744e

Observation a73daabc-cea4-4e13-a09e-18dec93b933a · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.325121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.325121Z digest=sha256:fd82418e8c2e3d2dfbe5821a7c1c21f201432f3e4ef919674c080e68cb89638d

Observation 3230d9c1-e31f-46e9-b892-1fe14743546a · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emu3: Next-Token Prediction is All You Need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.330544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.330544Z digest=sha256:3b5a65f9e786ae038ef17b7c55281c2a9d54fb856201f4596c55f01148976a22

Observation c4351010-997d-4a3b-badc-15655ea7642b · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.336002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.336002Z digest=sha256:283c9b7efa65e0b9dfb38c38e833c7cd78aab9f1815e18629c671412b94a4293

Observation 5335c9a1-029e-479f-bd32-a2e70ac05448 · outbound

This paper cites Advances in neural information processing systems35, 24824–24837 (2022).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in neural information processing systems35, 24824–24837 (2022)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.341793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.341793Z digest=sha256:a3545819779e2231a7d05b973527e13906542355d416ff5fc7712c3c4145553e

Observation 8fd01c75-0378-4401-9204-80ee2baefb7c · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.349765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.349765Z digest=sha256:4526f06025d74e22ebe9cd431c9670cb5f29116fb872f97b3d11eccf5d89c0e0

Observation be61ba36-fb07-4958-b08c-b7afd65ebeba · outbound

This paper cites OmniGen: Unified Image Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO OmniGen: Unified Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.355229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.355229Z digest=sha256:cc5f159dcdbac2c41a46b6c70b63299ffa1264319d9f87edb56d715a5bd78e03

Observation 53886b6d-17f0-4ff6-9054-18d186fb1d29 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.361913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.361913Z digest=sha256:34a6f6d8f97ed5f5776a88dddd749677e0810743a34aa93e449466143c654261

Observation ebd3e111-3b5e-4f16-a112-97f3ec83ebfa · outbound

This paper cites Advances in Neural Information Processing Systems37, 75329–75354 (2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems37, 75329–75354 (2024)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.367010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.367010Z digest=sha256:7ca0ff5a8336dff0527d924fd7883f3148e05856a0a0ce59a9c8ec323987b658

Observation 1e8245d0-d674-457d-bc56-822780fdfde5 · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.372378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.372378Z digest=sha256:7845cb98924faa510cba4842d43f6661b969825f5d83897e8aba7462d3ec577f

Observation 3b46e205-a7f2-4b77-bdad-2522772c0a0f · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.378298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.378298Z digest=sha256:a50bec4bda99b76cf3b9953920af5c78039756cdcc5b12c0814931081d047a4c

Observation 0d036f42-8e5b-4d3f-914f-31e1b6e806bf · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.383514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.383514Z digest=sha256:473443540de5060624417063eeabbb71f949617421290d2a3152493c54c4cd24

Observation caa0872c-e1fe-4ea6-a846-d49e2085b19a · outbound

This paper cites In: Forty-first International Conference on Machine Learning (2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Forty-first International Conference on Machine Learning (2024)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.222106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:26:58.389387Z digest=sha256:f491695128de3cce3617e4e94216b46de7c067ea39b1ccccf1344d0f4004eddc

Observation f0d08762-d7eb-44f3-9d89-59bb95314e85 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO MMBench: Is Your Multi-modal Model an All-around Player?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.394455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.394455Z digest=sha256:3a19723311aed4008d7515f3300d660ac0373a7181cf5cc5456e01803a4e65ea

Observation 343a32e8-aecf-439f-acaa-75b61d7c7461 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.399332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.399332Z digest=sha256:25b5549e1ea9c2f1d56fc2736d1c46d0ab299363658216e8692bc5afa3ecdb27

Observation 2da049e6-d79a-4250-93a2-0ed83facdd20 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.404285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.404285Z digest=sha256:2930af79508bd6c0587286700eb4d99ee1bf55129fd5a3b82f03cee698e132ec

Observation aafa572c-8539-4cdd-bbf2-ea8757a1f41e · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.409542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.409542Z digest=sha256:55b578aa6e8f964db1bf3bd4b6df1d4764a686f69552953ea24611d70cd2cd40

Pith citing papers

Observation 6bdc7cd4-ff7c-49a9-905d-1c637f54ce25 · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:10.990159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:10.990159Z digest=sha256:f5919bacccdb4172098299faec9ba0464ad2a43d6c5cbff269c1322271c98b98

Observation 1b66d188-aedf-42e4-9b63-f192e703d310 · inbound

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation cites this paper.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.624099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.624099Z digest=sha256:2e5e6b7fb5c195aaa852240a748b04a5745261657efd780100d32e7a5952ee6c

Observation 15fdf608-50a4-443a-b870-878a7cd630e2 · inbound

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation cites this paper.

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:26.605018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:26.605018Z digest=sha256:f2e0e63f776c33106ff9d0680459ec4fab77d5f064aaaa456cd581a49cc13f5e

Observation 21e7eaad-27fb-40a0-a9c9-acdb4d3b6e1d · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:d07cb39f91e1f9ce59198c9e1a2d600053b3278dae8187130d427725fa02f8e5

Observation a0967b82-5fb2-430d-ae00-a9f2fe1aabca · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.137676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.137676Z digest=sha256:159b80125462ceb057adb6f39d026bd3eed158cda378dad703976d17b15a32d1

Observation 17c4f5ff-d6f3-4ac4-b784-185498c4a051 · inbound

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing cites this paper.

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:00:50.371078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:23:50.614589Z digest=sha256:c57f31697f5a58c8a59388987787fd0bf0442a37ead8b4f03a640898b32d2594

Observation d7b75b5d-a3ce-440c-b68a-90040d96e677 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:43.694718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T17:57:08.606559Z digest=sha256:7756b7611c0fce3c8907acd551bac620d15f930c9c64dcc24ab64cd6e0d0659e

Observation f35ec559-82e5-404d-9f3a-262d9319b7cb · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.751849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:78c9bfde003f4fb869f05feab905692daca99c0ef33f0e37873adbf155f10a74

Observation 3c0041f6-f342-4a24-b0d8-2bfd251871a0 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:59.889877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:3e25a18097f82ac71d4565756583097ebf257906aa066d359ee18662bde94abe

Observation a6585bc5-4f90-46ac-91cc-5b059183f3b2 · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.442237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:586f04a6ea64a77b57815bac920bc9c630c81429305037c511ec2476be5aa7a9