Pith. sign in

Paper Citation Record · LEDGER

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 10 inbound Pith citation observations for arXiv:2505.13031.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13031 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:26:58.409542Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:10.990159Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:06:29.440180Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7431ca4d-00fc-4ac9-ba5b-c7a196d5b8f2 · outbound

This paper cites Qwen2.5-VL Technical Report.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.092350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.092350Z digest=sha256:1c4833e51aef3f49c7723e1ef7fa32f0bd2edb78bb4dd47c9676007deff52f6f

Observation a51e9c89-782f-4e8b-966d-59c6fb7bd70c · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.099262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.099262Z digest=sha256:01570ad5d6a94ac45468feb449ebfea2d0ed33f232ad4806bf5fa598ce7e5c45

Observation 389c266b-d5d5-4b08-8564-f988b4986220 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.105088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.105088Z digest=sha256:c88723f81581b6aca2d35f70d36c4718af38c60be1b573382a40b7ab8f2275e4

Observation 1b0cfad7-a596-43a5-9889-8a122e251c33 · outbound

This paper cites r1-v: Reinforcing super generalization ability in vision-language models with less than 3.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO r1-v: Reinforcing super generalization ability in vision-language models with less than 3

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.518115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.110901Z digest=sha256:c5f7c39ade60426069afaa8d2fdc4ec56f502f57b4c24d926866d42db2876e3c

Observation 8d82be19-ace8-4e34-a7fa-675d7a3ffa2d · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.116559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.116559Z digest=sha256:74f4d3cb3da2efcf73a76a685a300ff8e739e148583c95d598af1aa64e52909b

Observation 766af505-288d-4f51-95e5-2adbc908af8d · outbound

This paper cites Science China Information Sciences67(12), 220101 (2024) 10.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Science China Information Sciences67(12), 220101 (2024) 10

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.488898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.121570Z digest=sha256:19dc4cfba2e9dad4bf5f954c716a3b40d6f09c8ea2d6f1c7636228c07607aeaf

Observation d9523808-2f9f-45e0-8845-8718a37b2ac6 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.126897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.126897Z digest=sha256:d447808f71d5f5541127ae112d4263515c31d357099cc08421bcb77d275c37d2

Observation 83349dd2-4b7d-4060-8d87-1c24a8796109 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.133735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.133735Z digest=sha256:d2a123c2743bdb85b39bee094b7944bf83ee9eb09311b705dfa1ee8261eda6ab

Observation 71cab7d0-ea13-4b25-88de-9852fedef3cd · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.139315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.139315Z digest=sha256:55e798f821b19cdf9e3a8c98cd45fc43abf9f9472b9fd960edccc4e586e7251c

Observation eee365db-c203-4bf0-a66f-19a157c902cb · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.146231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.146231Z digest=sha256:fdb64dd0da3692ffbeb53da2de4bf00019b023d9f7460a3a30ead61aafa8e8f6

Observation 257300d6-e990-4252-a71f-e6d7c8cfda74 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.153083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.153083Z digest=sha256:18c818c8d8b3e9d42fb8573f3d11502e174742a877d998afb1ef19c294111e01

Observation e2edbfb5-4dce-4b66-9659-d94dd2214f9d · outbound

This paper cites Advances in Neural Information Processing Systems pp.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems pp

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.470305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.158664Z digest=sha256:7d702780b35628462bde3ec41675a48ca0f676175fe1f83f79c8ffc3833617b8

Observation 25e2075e-38a0-460a-90fe-77141058ab01 · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:26:59.450727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.163671Z digest=sha256:e298da707ea05196acfdb2ed97cc1320eba32ad4ee2f74e81368898b1215848b

Observation 3fcd54f2-d329-46df-825b-d4fff51a8731 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.170393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.170393Z digest=sha256:b0cc9f858f8853d5bbd08bfe59b74cbe6568409e9e7d0748283de3aadc0dccc0

Observation 7528f7ae-75b6-40a4-bb11-02e6e231d17d · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.175229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.175229Z digest=sha256:6622ddd83bd100cc97a3ac64a9afdcebd7327c6c790de62fcb5a5128efddfbb1

Observation daf4f96a-731d-42e2-8f39-e4b1da2c1424 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.180483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.180483Z digest=sha256:4a5dd3769492813ef2eb8776598dd9565964fb105f67340cf8dada0f8fb16c5d

Observation a90a0824-2081-47b8-928f-21aa3cabc3d1 · outbound

This paper cites ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.185469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.185469Z digest=sha256:a8b374fe4f91d5363638fbd226a10e2bd528df2120b9544ef455d188bd66de40

Observation 8f7ffe76-90bb-44b7-ad54-10d5e720510a · outbound

This paper cites ARGS: Alignment as Reward-Guided Search.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ARGS: Alignment as Reward-Guided Search

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.192292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.192292Z digest=sha256:773e9a4eba9ec29c8bbc4fd3936f45c48ec3b885cbbb9b098f1b961493838b24

Observation a4882627-c6d9-4a15-b758-8c0764825b81 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Training Language Models to Self-Correct via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.197909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.197909Z digest=sha256:5df981cc485cdcfcfee655ca97c06272d3f88db87625e8826edd8a635985567f

Observation 970508f6-7142-4887-a75f-fbc7eeecfd46 · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.203015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.203015Z digest=sha256:244ec8284203db073c6eb395d09c98d0e25607e092fd36357f00d36e7baea989

Observation 18aca474-96d7-4523-ba18-750605542886 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.208127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.208127Z digest=sha256:e3f4305adbe7aa80c20d575d855723140ccacbb24eca67cbc09bba9a1d241d71

Observation 6a3dfc0b-2b32-446e-a4c0-85130968bbd0 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.214608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.214608Z digest=sha256:e71f24ff647bb702c4fec27607b2c2a928e5225395d80587ad422a28ae824a3c

Observation 97f83c26-75db-462e-ab39-20c7d35cdd82 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.220542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.220542Z digest=sha256:e4d72eb917c1b18d77f39661d93202997b3684d486214418c586ddc15ebe8f9e

Observation 7598a37a-d303-4a7a-b95d-38b2f95001f9 · outbound

This paper cites https://openai.com/o1 (2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO https://openai.com/o1 (2024)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.420022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.225798Z digest=sha256:63573f54e2bd4f6fdd830981fe8fe3d0aa7b03f734ce5ebbce9fda60b26219f9

Observation 1a45b6d6-9c64-4ef2-be54-7642c52af2ef · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:26:59.401571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.231910Z digest=sha256:ea0093e5f617bf657b7280436111af44bf21526e0143f4fb3f5f12b79d5b3245

Observation 291b28c4-8111-430b-9d8c-04cdec535f42 · outbound

This paper cites Transfer between Modalities with MetaQueries.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Transfer between Modalities with MetaQueries

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.237599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.237599Z digest=sha256:7309a26a3477c8eb33be83cb9ee365c8d9858d2a0b6892290fe36ea3d9aa1cd7

Observation 85580cca-6308-4b8d-8487-64a66c2ff527 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.243052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.243052Z digest=sha256:5a20353c9c99c3b309fbd8380859a85e682c18efa1eb6e18a6d9fbc1630edc6c

Observation 919c26c5-2940-41b1-b90f-853ae5be2a05 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.249057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.249057Z digest=sha256:e5d3ca638fdb3d36aefcf07f230507abd20dd7cf14b65653741f3ddfbf64afa0

Observation c2050e8c-61c5-498d-b050-32f0d785b733 · outbound

This paper cites Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.255579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.255579Z digest=sha256:77dc8c545b3a6063bfde6c6f07c47489dc76e1b34299161fe97a70e22ec6b870

Observation ec67038c-8d8a-4cdd-b29e-5fc2732a4d94 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.261282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.261282Z digest=sha256:cf47aff11cb54c55f462948da3fa00444599b1e24dff8c54a7b92d742a207f9f

Observation 28a7daf4-7488-4264-a553-191935772155 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.267389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.267389Z digest=sha256:633001bb0c1875fc94e969d2798294828ba527eac96ded372a7b2f31bd46ca57

Observation fd21ddc3-1885-4e10-86b5-078be8500a1c · outbound

This paper cites URL https://laion.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO URL https://laion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.361428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.272917Z digest=sha256:c4e42d7b7c6b79087d77f274647dac901bc8b608ab4c1c5c88bbb8fb15c67f88

Observation dc57ee99-1fd9-474c-8474-4f5352a5cc1e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.278116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.278116Z digest=sha256:163ff4ebcefd24de159f8f94172c9deb634e90179044f2083680a206f4b41987

Observation 201000fc-ad77-4f4e-bc67-e3e3d979dec3 · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:26:59.340085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.283274Z digest=sha256:50ed56150454090df3f9c98c42eb7fa164451c14a2fdf43edaee91296903cd27

Observation 3c262794-8a07-4f9e-bede-c27c41c66ddf · outbound

This paper cites Advances in Neural Information Processing Systems36(2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems36(2024)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.288057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.288057Z digest=sha256:328397c2283e95995be7f231b0ed162f07173c8401eb6d8656db5b3d6ba23cd9

Observation 4c803fed-c3c5-4f6d-8871-30da9ed25ac7 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.292902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.292902Z digest=sha256:287381e282a6f84622876ac533be355468a53228c96f75ee93f7bb0d52ffb081

Observation e4632963-3123-4104-b269-d512bd4d163d · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.298120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.298120Z digest=sha256:01812a0413e5d2c3e9ec7b52db7952d7ec23c2f52a8cb4ae64c514d18c7fb0e3

Observation 98a15ab1-62d7-4d53-9d89-e417288b8476 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emu: Generative Pretraining in Multimodality

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.303151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.303151Z digest=sha256:48bc5eef786bfa0e0673af59c53609d08a235bdf6a4f5f8acc7c730bc6921ccc

Observation c29421bb-db06-4785-8940-54aa1b3a4209 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.308208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.308208Z digest=sha256:8cff28397ccb67c797ecfd4a3b6005db970c625e4170212d4bed61e13f955074

Observation 66b03bf8-fcd7-4704-bded-6c384a40c4e6 · outbound

This paper cites https://qwenlm.github.io/blog/qwen3/ (2025).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO https://qwenlm.github.io/blog/qwen3/ (2025)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.291607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.313105Z digest=sha256:6769b667a569865898c41ca479ec55dbe1558902faf19b577ace9aecc782df71

Observation 24f5c3f4-324c-4a48-9a70-8ebbefc79636 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.318304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.318304Z digest=sha256:d70913a3621b5f7580faf1d646e5e1e17d8654e331f86897886ba870fa4cd101

Observation a73daabc-cea4-4e13-a09e-18dec93b933a · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.325121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.325121Z digest=sha256:5549adad4d555f7df2b8ee01e76a70e0db5a2ea7cb65e1ece9ea8cdf6b66bb63

Observation 3230d9c1-e31f-46e9-b892-1fe14743546a · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emu3: Next-Token Prediction is All You Need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.330544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.330544Z digest=sha256:e8ddc747ac020aac3872727fb4ac6528e72a755c9846f2d12908479c3507a3bb

Observation c4351010-997d-4a3b-badc-15655ea7642b · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.336002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.336002Z digest=sha256:e05129747b0134ca2a1e5a174a3ae2c8dac1580e9db2af5c5fade172efac34b9

Observation 5335c9a1-029e-479f-bd32-a2e70ac05448 · outbound

This paper cites Advances in neural information processing systems35, 24824–24837 (2022).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in neural information processing systems35, 24824–24837 (2022)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.341793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.341793Z digest=sha256:4045c00939ad2232c8a4ece303e36fad777af18525482c2c55c54d04f2235175

Observation 8fd01c75-0378-4401-9204-80ee2baefb7c · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.349765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.349765Z digest=sha256:621fce1e6cdcebad6b3146d6162d088da6120318bc9f0df0d8542d9e0748b5db

Observation be61ba36-fb07-4958-b08c-b7afd65ebeba · outbound

This paper cites OmniGen: Unified Image Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO OmniGen: Unified Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.355229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.355229Z digest=sha256:f2398591d7dc6912746b7324ce8bb2c9e5057932fdecc9c1e4b7dd8d41639cbb

Observation 53886b6d-17f0-4ff6-9054-18d186fb1d29 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.361913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.361913Z digest=sha256:a822d6a56ab50026a4b0da823c6f7039d8dd13fb476f93e958719d8ad2628d21

Observation ebd3e111-3b5e-4f16-a112-97f3ec83ebfa · outbound

This paper cites Advances in Neural Information Processing Systems37, 75329–75354 (2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems37, 75329–75354 (2024)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.367010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.367010Z digest=sha256:17c778367e112a8b37acacd14d6bb1ffacfc0447b2731f83b214cf343d0c54f8

Observation 1e8245d0-d674-457d-bc56-822780fdfde5 · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.372378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.372378Z digest=sha256:9bfe9bc94f6cea9a8aa52d82e30c4adee46fd4be552b9f88bda9016beef2a843

Observation 3b46e205-a7f2-4b77-bdad-2522772c0a0f · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.378298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.378298Z digest=sha256:af84caeea87723eb771afe11cc2b941978b6ed08640c735134b0473e3f9b15b6

Observation 0d036f42-8e5b-4d3f-914f-31e1b6e806bf · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.383514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.383514Z digest=sha256:618deb44ecdfe52f9f85824c6c5cc0ec7105eabfc183d2bdf82225086c05b8e5

Observation caa0872c-e1fe-4ea6-a846-d49e2085b19a · outbound

This paper cites In: Forty-first International Conference on Machine Learning (2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Forty-first International Conference on Machine Learning (2024)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.222106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.389387Z digest=sha256:5e5cf68d1b47d28716d50abbd9edfd5f222a75ba5201a05e32cc3cfdbfa847a2

Observation f0d08762-d7eb-44f3-9d89-59bb95314e85 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO MMBench: Is Your Multi-modal Model an All-around Player?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.394455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.394455Z digest=sha256:9fcb6a54c9eb90ba85ab94408b09276f0f12915ad3dafb2fa78556b677f40602

Observation 343a32e8-aecf-439f-acaa-75b61d7c7461 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.399332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.399332Z digest=sha256:9f1aa16f1df4596946750e7d8de37b9450eff1a4b1070b715ff010588f244060

Observation 2da049e6-d79a-4250-93a2-0ed83facdd20 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.404285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.404285Z digest=sha256:3f77e82cf8da819ce00ecc5ddbbcb8563242262df1ad71898567001b3f387e16

Observation aafa572c-8539-4cdd-bbf2-ea8757a1f41e · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.409542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.409542Z digest=sha256:39d4f3f6bdab2bdb555e5c2d3b1352ab61cea61a4dadae7b90c1accd41855be2

Pith citing papers

Observation 6bdc7cd4-ff7c-49a9-905d-1c637f54ce25 · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:10.990159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:10.990159Z digest=sha256:b77c6d163d843aeb82b1c245c03a36c146094a4911813f6778faf8aebb58e0a5

Observation 1b66d188-aedf-42e4-9b63-f192e703d310 · inbound

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation cites this paper.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.624099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.624099Z digest=sha256:5ee8b0c6a89d343a867fc2f9746f662034ec15b43626473bd627f5d74143000a

Observation 15fdf608-50a4-443a-b870-878a7cd630e2 · inbound

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation cites this paper.

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:26.605018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:26.605018Z digest=sha256:5b99863183be93e6a2746c2d724180923210eae40b8cc55ab0e375ddd138f86d

Observation 21e7eaad-27fb-40a0-a9c9-acdb4d3b6e1d · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:e8b4bce9c4803b5b240d3afd1a6d83f2e9216e1a7e957b4e22940b8fb06d1e7b

Observation a0967b82-5fb2-430d-ae00-a9f2fe1aabca · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.137676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.137676Z digest=sha256:4510ba6f59cb570531b36ee97a36bc49fb251a34ab31135ff2d86ca9b8cb5cef

Observation 17c4f5ff-d6f3-4ac4-b784-185498c4a051 · inbound

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing cites this paper.

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:00:50.371078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:23:50.614589Z digest=sha256:0321d350c70b3aa3254309eed7614ec4dc81572b7abf5442a16aa38734096535

Observation d7b75b5d-a3ce-440c-b68a-90040d96e677 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:43.694718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T17:57:08.606559Z digest=sha256:7595c5b7cfbc3581900fc2c8408b759ff7327897d8aa8952e7e41999cd42678e

Observation f35ec559-82e5-404d-9f3a-262d9319b7cb · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.751849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:ae4a24301d9c7d2f837630a706b7714f1cfad54cec49bdc28271ac57ffd386c6

Observation 3c0041f6-f342-4a24-b0d8-2bfd251871a0 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:59.889877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:da7e2a32b67c6a210da62943f088f9b9a9b9b89d5402fb5d4daaa85631f31980

Observation a6585bc5-4f90-46ac-91cc-5b059183f3b2 · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.442237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:6615c4514716470c343746b0979d6416b06c57136a0b8d2c37dff0c92a59fee7